JSON plus comments and python-style multi-line strings is great.
The thing where you can leave quotes off strings makes me nervous, especially the example where the value is HTML with its own embedded double quotes for attribute values.
Not requiring quotes on strings like that looks like an obvious vector for injection attacks. I guess Hjson isn't designed to be generated automatically, but I'd prefer a format that is easy to generate safely.
What I really want is JSON plus comments plus multi-line strings plus relaxed rules on trailing commas... While maintaining as simple and unambiguous a parsing model as possible.
While I have been hoping for "JSON plus comments" to be a real and common thing for quite a while now, one of the strengths of JSON is right there on json.org. See it? A set of five simple syntax diagrams that entirely and virtually unambiguously define the json syntax.
It’s tough to know when to stop when simplifying a syntax. (For an extreme example, see Stylus, which was like Sass but so extreme that mixins and properties became ambiguous for each other.) I, too, would like to see the return of quotes for string values, for increased clarity.
> When you omit quotes the string ends at the newline.
> Preceding and trailing whitespace is ignored as are escapes.
>
> A value that is a number, true, false or null in JSON is
> parsed as a value. E.g. 3 is a valid number while 3 times is
> a string.
It is ambiguous only if you think it can be a completely useless format. If you decide to use it, chances are you considered it useful, and the only useful way this can parse is as a number, otherwise it would't be able to represent numbers, booleans and null. I'd say it parses pretty much as expected, and as such, contrary to other comments, you don't have to go back to the docs every so often.
The problem is that this is supposed to be a more-easily-readable form of JSON, but it requires consulting the docs to understand the meaning of something which is completely unambiguous in regular JSON.
I came here to make the same comment. The ability to leave off quotes on strings is a misfeature which overly complicates the language. Also I see no need for three different comment styles.
Yes! There is the same ambiguity in Yaml, so syntax highlighters of the format don't agree where a string ends...
The site even has an ambiguous example:
three: 4 # oops
Well, is that "4 # oops" or 4?
I saw no rule about ending comments, and they say that
three: 3 times
is a string.
I have seen no formal specification of the grammar, so we already have lot of ambiguity. Good luck to implementers...
Seriously, removing the "no quotes needed" rule would improve greatly the format. If you want to include HTML with double quotes literally, just use the multiline string format and be done.
That might work for programming languages (C++'s bad rep nonwithstanding), but for data exchange formats you cannot cherrypick features, you have to support the whole spec, and nothing else.
I think it's more nuanced than that. You must be strict in what you emit, but can be liberal in what you accept. That liberalness can go too far though, and should not make parsing brittle, or encourage misuse of the spec. It's there more to allow someone's unambiguous mistakes to still parse.
- Accidentally left in a final comma on a list? That's okay, that only means one thing, we understand.
- Allow non-quoted keys on objects? Well, we understand JavaScript generally allows this, so we'll let it slide. This time.
- Make newline significant and define new items? Okay, are we just ignoring space efficient payloads now? Should making it space efficient mean changing formats from Hjson to json?
- Considering all terms in place of a object value a string until a newline? Are you just trolling me now? How is that more human readable? Does your spoken language not use quotes to distinguish distinct chunks of communication or something, and if so, does it use a Latin alphabet so it's off-putting when you see them?
Needless to say, I'm really confused by the reason this even exists.
I think the problem it is solving is that JSON is designed and best used as a data exchange format, but it also gets used for configuration files, which it does okay with but is not really so good. INI files don't have a clear standard. YAML is too complicated, and using turing complete javascript for configuration seems like you've just gone too far.
we just need JSON, but with a couple things fixed up to make it nicer to use for configuration files.
> we just need JSON, but with a couple things fixed up to make it nicer to use for configuration files.
Using JSON for configuration is just the whole situation of using XML for data exchange redux. One of the major points for JSON over XML for data exchange was that it was so much better because it was optimized for data, not markup. Why are we ignoring this argument now that JSON is on the other side? JSON is used for configuration because it's ubiquitous, not because it fits the problem domain well. Let's just choose a more appropriate format.
Choosing the most common set of rules for INI files (what is proposed by Wikipedia[1] is probably sufficient) would serve us MUCH better than coaxing a data interchange format into that role.
If you're using a strongly typed language, XML even has one (IMO) massive advantage over JSON: You can use XSD to define a schema declaratively. This means you get
(1) lots of general tooling support, in particular you get at least decent editor support for your config file, and
(2) you can autogenerate the code needed to read your configuration into structured data without having to do any unnecessary duplication of "key names" (tags) as you have to with e.g. INI or JSON. You also get the data sturctures themselves for "free" (based on the XSD).
Alright, it's not the end of the world to not have these things, but they're both very nice to have.
Yes indeed I have, and I have several problems with it, one of which is "Expires: August 3, 2013" with no new version in sight. My other major problem is that AFAICT it doesn't support one of my pet favorite features namely "Algebraic Data Types"[0, 1]. If you extend json-schema like Swagger has done it might support ADTs properly AFAICT, but I don't have any actual practical experience with Swagger.
[0] https://en.wikipedia.org/wiki/Algebraic_data_type
[1] I should note that support for ADTs in XSD is sometimes sketchy in the various code generators, but at least the XSD specification supports it. (It could also be argued that XSD supports something even more general... which it probably shouldn't since ADTs basically cover the whole data structure space unless you go to higher kinds, inheritance and such.)
well basically virtually anything other than JSON that's actually designed to be configuration would be better for configuration. There's dozens of them. that is the problem: JSON parsers and generators are ubiquitous in the way that no single configuration format is. Right now, I can just use json in any language and get data between any language and any other language. If I use it for config I get the advantage of even being able to config across multiple languages that might be getting used in a single system if I have to. No proper configuration format has that level of mindshare and interoperation.
While I agree on some points, I do not agree in general. I think too much prescriptivism in protocol implementation is a naive approach, and assumes that we can always get things right initially. Sometimes real-world concerns and needs drive changes, not just sloppiness.
Well, the first thing I thought of is what nightmare would it be to safely implement a parser for this in C. I filed a Github issue for this one: https://github.com/laktak/hjson/issues/37
i've used json5 since the beginning. it's great! just a couple week ago they merged in descriptive error messages too. No more wondering where the bug is!
Strings without quotes leads to all kinds of trouble in YAML. You end up just quoting everything anyway the first time you need to use "false" or "[some words here]" as a string.
> The thing where you can leave quotes off strings makes me nervous, especially the example where the value is HTML with its own embedded double quotes for attribute values.
Learn from Perl. The quote operator is your friend (and I frequently lament it's omission in Bash). You could simplify it by not using the matching enclusures ({ and }, [ and ], etc). It's easy to parse. and if you keep the quoting character somewhat rare, it's not hard to read.
I haven't used Perl in quite some time, and this is the sort of thing is why it was bad. Quotes are quotes are quotes in almost every language. It's completely unambiguous, the downside is that you sometimes need to escape them.
This, on the other hand, is a 'solution' to escaping quotes that is completely mad. Using non-standard quotes, especially mixing and matching them is a disaster for readability and maintainability (using a T in your string now? need to change the quotes!). Triple quotes are just find if you want to avoid escapes, and hjson seems to support them.
> This, on the other hand, is a 'solution' to escaping quotes that is completely mad.
Meh, Perl's solution is fine. You can throw up your hands and say it's crazy, but as a person who worked with Perl for 20 years, I've never had the problem you describe. I tend not to use the qq() or q() style quotes, but I've used s@@@ and s,,, so many times I can't count. It's really quite nice (and perfectly readable unless you do something weird like 'sxxx'.
> Quotes are quotes are quotes in almost every language. It's completely unambiguous
Oh, like in C and C++ where single quotes denote a char, and double quotes a string?
Or in Perl, PHP and Ruby where double quotes interpolate, and single quotes don't?
Or in JavaScript and Python, where there's no functional difference between single and double quotes?
Or C#'s string literals which only support double quotes, but you prefix the string with @ to denote it's verbatim?
Or systems that allow repeated double quotes within a double quote string literal to stand in for an escape ("foo ""bar"" baz"), as many SQL systems do?
Or what about systems that interpolate, and the differences between what they do and do not interpolate? Variables? Escape characters? Hexadecimal escapes?
You're fooling yourself if you think it unambiguous in anything except for the language you are dealing with, and if you're within that language, who cares what you use as long as it's consistent? You learn it, and then it's unambiguous (if implemented well).
This is no different than if your language supports hex numbers (usually done through prefixing it with 0x). Those are two different ways to specify the exact same thing (a binary number!). The benefit comes from using it in the circumstance where it's appropriate. That is, where it enhances readability, not where it detracts from it.
> This, on the other hand, is a 'solution' to escaping quotes that is completely mad. Using non-standard quotes, especially mixing and matching them is a disaster for readability and maintainability
because I think it's clearer, and learning once that a literal q defines a new quote operator that is in effect until it's next seen is simple, easy to remember, and yields very useful readability gains.
> using a T in your string now? need to change the quotes!
I included the qTT example just to show how it worked, not to endorse its use. I thought that would have been obvious from my statement "You could simplify it by not using the matching enclosures ({ and }, [ and ], etc). It's easy to parse. and if you keep the quoting character somewhat rare, it's not hard to read."
In any case, I fail to see how how that's a problem beyond any other quote character. Including that character in your string will result in a compile time error in all but the most esoteric of cases, making it easy to find.
Do checkout http://json5.org/ based off the original JSON author Doug Crockford's own proposed parser extensions, primarily trailing commas and comments, both of which have been a point of contention pretty much the day since the day JSON landed. It's a much simpler and saner proposal.
Trailing commas and JSON comments are are already supported in the newer browsers (try the Chrome console for instance).
Fortunately quoteless strings or optional-commas/newline-separator as proposed in Hjson will never fly. They are brittle and ambiguous. Who knows what will this get parsed as:
{
a: hello's and hi's have
'misplaced' apostrophes
b: ball: a round # and # bouncy object
c: cakes and
candy: both have sugar
# but how do I include a hash at the start of a multiline-unquoted string?
}
Javascript object literals != JSON. JSON is a restricted subset of JS object literals (and not actually a strict subset: a JSON string can contain unescaped U+2028 "LINE SEPARATOR" and U+2029 "PARAGRAPH SEPARATOR" codepoints, a Javascript string can not)
YAML is more complex that most people tend to realize. (This was brought up in a 2011 discussion about possibly standardizing a metadata section for Markdown documents which sadly went nowhere. [1])
Take a look at example 2.11 in the YAML spec [2], for example, and see if you can make heads or tails of it.
You don't need most of those features. A pared down YAML with the cruft removed (implicit typing, flow style, tag tokens, node anchor & references) is actually pretty simple as well as less "gotcha-y".
Okay, does anyone actually believe that causes a problem?
I can think of ONE time when that causes a problem, and that's with indentation with multi-line strings. Oh look, HJSON included that feature. That's like throwing the baby out and keeping all the bathwater.
I can't find a specific example off the top of my head but I'll say I've been managing a Jekyll site for a while now and whitespace errors in frontmatter and data files cause all kinds of problems. I'm not sure I could explain the details but it's a legit criticism of YAML. IMO part of the problem is that YAML looks very straightfowrard and is until it suddenly isn't. Whitespace is part of that problem.
As an anecdote on the flip side, I've been building Middleman sites for a while now and can't remember ever having an issue with whitespace in the front matter or local data.
Unquoted strings are valid in yaml just like this format. There are at least 2 ways to specify a list of things. There are some super bizarre looking possible formats for lists of mapping types.
There are a number of others given the length of the spec. Yaml is a complicated beast that generally has more than one way to do any given thing
I like that you can leave the quotes off of keys, which should always parse as identifiers. (And those that don't should require quotes.) Leaving the quotes of values seems like the problem.
Primary goals were to remove as much syntax as possible and make it play well with line-based diffs (with the hopes that someone who knows knowing about the language could resolve conflicts without getting tripped up by surrounding quotes, trailing comments, etc).
Unfortunately, conflicts in white-space based languages can get even worse than regular conflicts because you have very few visual structural "anchors" to start to gain an understanding of the conflict. (If you have to resolve manually, that is.)
Granted, if the number of conflicts which cannot be automatically resolved is reduced by enough, then it might not matter in the grand scheme of things. However, I'd be worried that this would make "accidental" automatic resolution of semantic conflicts more common. That may be an unfounded/irrational fear, I don't know.
The thing where you can leave quotes off strings makes me nervous, especially the example where the value is HTML with its own embedded double quotes for attribute values.
Not requiring quotes on strings like that looks like an obvious vector for injection attacks. I guess Hjson isn't designed to be generated automatically, but I'd prefer a format that is easy to generate safely.
What I really want is JSON plus comments plus multi-line strings plus relaxed rules on trailing commas... While maintaining as simple and unambiguous a parsing model as possible.