JSON

JavaScript Object Notation, a lightweight data interchange format that is easy for both humans and machines to read.

JSON (JavaScript Object Notation) is a lightweight, text-based data interchange format that represents data as key-value pairs. Douglas Crockford put the notation together in 2001; it was written up as a formal standard only later. The first standardization was RFC 4627 in 2006, followed by the first edition of ECMA-404 in 2013 and, in 2017, RFC 8259 together with the second edition of ECMA-404. As of 2026, the two documents to consult are RFC 8259 and the second edition of ECMA-404; they are worded differently but define the same grammar. Although derived from JavaScript syntax, JSON is a language-independent, general-purpose format supported by virtually every programming language, including Python, Java, Go, and Ruby.

JSON supports six data types: strings (enclosed in double quotes), numbers (integers and floating-point), booleans (true/false), null, arrays (square brackets), and objects (curly braces). This simplicity is JSON's greatest strength and the reason it became the de facto standard response format for REST APIs.

Compared with XML, JSON needs no opening and closing tags, so it is less verbose and faster to parse. The same data comes out shorter than in XML, where every key name is written twice, once in the start tag and once in the end tag. The actual reduction, however, varies widely with the length of the key names, the depth of nesting, and whether attributes are used, so it is not a fixed ratio worth memorizing. If size is going to drive the decision, write the real data out in both formats and compare the byte counts after gzip. For data in which the same structure repeats, gzip absorbs the duplicated tag names, so the gap after compression is smaller than before it. XML, on the other hand, has mature schema definition (XSD) and namespace support and remains the choice where strict data validation is required. YAML version 1.2 was designed to take in JSON as a strict superset, so any JSON text can be read as YAML as it is. Because YAML has comments and anchors, it is favored for configuration files, but its indentation-dependent syntax is prone to errors when copied and pasted. Note that the containment is not a perfect equivalence: a JSON document with duplicate keys, for example, is invalid under the YAML specification (some implementations report an error, others silently keep only one of the values).

JSON has several limitations. It does not support comments, so JSON5 or JSONC (JSON with Comments) is sometimes used for configuration files. There is no date type, so date and time values are conventionally represented as ISO 8601 strings (for example "2025-01-15T09:30:00Z"). If the trailing Z or offset is left off such a string, the result shifts depending on whether the receiver interprets it as local time or as UTC, so always include the time zone. Trailing commas are not allowed either: a comma after the last element of an array or object is a syntax error.

Although the grammar itself is simple, the specification leaves details to implementations, and in practice two points cause most of the trouble. The first is numeric precision. RFC 8259 does not specify how many digits or what range a number may have, and most implementations assume IEEE 754 double precision (binary64). The same RFC gives 2 to the 53rd power minus 1, that is 9,007,199,254,740,991, as the largest integer that interoperates reliably (the same absolute value on the negative side). Passing an ID beyond that as a number rounds off the low-order digits and ends up pointing at a different record. A 19-digit ID or account number is safest exchanged as a string from the start. The second is duplicate keys. RFC 8259 only says that names within an object should be unique; it does not prohibit duplicates. What happens when they occur is left to the implementation: some keep only the last value, some raise an error, and some return all of them. Because this is also used as a technique to slip past validation, the producer must not create duplicates and the consumer must detect and reject them as well.

In practice, JSON Schema is widely used for validation and schema definition. Defining the structure of API requests and responses in JSON Schema makes the data contract between client and server explicit and keeps invalid data out.

From a security perspective, JSON from an untrusted source must never be parsed with eval(). Always use JSON.parse() to prevent execution of injected code. Also, if JSON is assembled by string concatenation, a double quote or curly brace in the input can rewrite the structure of the data itself. Rather than escaping values by hand, leave the embedding to the serializer of each language (JSON.stringify() in JavaScript).

From a character counting perspective, JSON syntax elements such as key names, curly braces, square brackets, double quotes, colons, and commas all affect data size. As for character encoding, RFC 8259 requires JSON exchanged outside a closed ecosystem to be encoded in UTF-8 and forbids adding a BOM at the start (a receiver may ignore one). Byte estimates can therefore assume UTF-8. Minification removes unnecessary whitespace and indentation, and combining it with gzip compression shrinks the transfer size considerably. In responses that contain Japanese text, the byte count also depends on whether characters are written as they are or escaped as a backslash followed by u and four hexadecimal digits (such as \u3042). A Japanese character is 3 bytes in UTF-8 but 6 bytes when escaped, so a library setting that mechanically escapes all non-ASCII characters roughly doubles the size of a response that is mostly Japanese. When optimizing API response size, dropping unnecessary fields and shortening key names are also effective.

Share this article