Appearance
Unicode syntax
Source: Elixir reference, Unicode syntax.
Strings ("olá") and charlists support Unicode since v1.0 and keep their contents exactly as written. Strings are UTF-8, charlists are lists of code points. Since v1.5, variables, atoms and calls also allow Unicode. String.Unicode.version() shows the Unicode version of your installation.
Introduction
Unicode characters must still obey the syntax rules: variables and calls can't start with an uppercase letter. The allowed characters are the ones Unicode specifies for the writing systems of languages still in use. Emoji, alternate numeric forms and musical notes are excluded. Many restrictions exist for security:
- Normalization. "josé" can be written as
j o s é(NFC) orj o s eplus a separate accent character (NFD). Elixir normalizes to NFC. - No mixed scripts unless separated by
_. A variableаdmin, with a Cyrillicаand Latin letters, raises** (SyntaxError) invalid mixed-script identifier found: аdminand lists each character's script. - Confusable identifiers in the same file warn:
а(Cyrillic) andа(Latin).
Unicode Standard Annex #31 (identifiers)
Elixir conforms to UAX #31 (version 17.0).
R1. Default identifiers: <Identifier> := <Start> <Continue>* <Ending>?.
<Start>: uppercase, lowercase, titlecase, modifier and other letters, letter numbers, plusOther_ID_Start, minusPattern_SyntaxandPattern_White_Space. In set notation[\p{L}\p{Nl}\p{Other_ID_Start}-\p{Pattern_Syntax}-\p{Pattern_White_Space}], normalized to NFC.<Continue>:ID_Startplus nonspacing and spacing combining marks, decimal numbers, connector punctuation andOther_ID_Continue, minus the same two sets.<Ending>: an Elixir addition, only?(003F) and!(0021).<Medial>is empty. ZWJ and ZWNJ are not allowed (R1a isn't implemented), and bidirectional control characters aren't supported.
Atoms follow that rule, with _ also allowed at the start and @ also allowed in <Continue>. Quoted atoms (:"hello elixir") allow anything, and all operators are atoms.
Variables and calls allow _ at the start and exclude uppercase (Lu) and titlecase (Lt) letters at the start: [\u{005F}\p{Ll}\p{Lm}\p{Lo}\p{Nl}\p{Other_ID_Start}-\p{Pattern_Syntax}-\p{Pattern_White_Space}].
Aliases allow only ASCII, starting uppercase, with no punctuation.
R3. Pattern_White_Space and Pattern_Syntax: Elixir supports only \t, \n, \r and space as whitespace, so it doesn't follow R3.
R4. Equivalent normalized identifiers: identifiers are case-sensitive, and atoms and variables are normalized to NFC (since v1.14). Quoted atoms and strings are not verified and keep their form.
Choosing R4 excludes R5, R6 and R7.
Unicode Technical Standard #39 (security)
Elixir conforms to UTS #39 (version 17.0).
- C1. General security profile: identifiers with code points in
Identifier_Status=Restrictedaren't tokenized (with the normalizations below). The often invisible HANGUL FILLER, for example, triggers a warning. - C2. Confusable detection: identifiers that look alike but differ, within one file, warn. Example: in
а = a = 1the two a's are Cyrillic and Latin, and in力 = カ = 1both are Japanese but different code points. Identifiers made only ofa-z,A-Z,0-9and_are exempt, because programmers already handlel/1andO/0with fonts. - C3. Mixed script detection: mixed-script identifiers aren't tokenized unless split by underscore chunks (
http_сервер).幻한is fine because Han mixes with Japanese and Korean (UTS 39 5.1), while Latin with Japanese needs an underscore::T_シャツ.if аdmin, do: :ok, else: :errfails with a descriptive error. - C4, C5: inapplicable. C4 (restriction levels) classifies arbitrary strings, and C5 (mixed numbers) doesn't apply since Unicode numbers aren't supported.
Additional normalizations
Since v1.14 some restricted code points are normalized to unrestricted ones. Currently only MICRO SIGN (µ) becomes Greek lowercase mu (μ). The normalized character gets the union of both scriptsets ({Greek, Common}), and Common intersects every scriptset, so it can appear in any script without causing mixing. The affected code points are ones the community uses, such as for microseconds.