Character Literals
Character literals represent Unicode scalar values (code points) and have type
char (types).
Use char for:
- single-character markers and delimiters (e.g.
',',':'), - working with code points when interfacing with parsing/lexing logic,
- representing control characters (
'\n','\t','\0').
If you need multiple characters, use string literals (literals string).
Notes#
What works end-to-end today (lexer → parser → checker → lowering → codegen):
- UTF-8 character literals like
'x','é', and'😀'(exactly one Unicode scalar, encoded in UTF-8 in the source file). - Escape sequences:
\n,\r,\t,\0\\,\',\"\xNN(exactly two hex digits)\u{...}(1–6 hex digits)- Equality and inequality comparisons (
==,!=) overcharvalues. charvalues are lowered as au32scalar in the current IR backend subset.
Not implemented yet (or not specified as stable):
- A dedicated diagnostic for invalid character literal spellings (most invalid forms currently surface as generic “unsupported expression” errors in the Supported forms).
Surface Syntax#
Character literals are delimited by single quotes:
let a: char = 'x';
Rules:
- The contents must represent exactly one Unicode scalar value.
- A character literal must not span multiple lines.
- The source file is interpreted as UTF-8.
Escapes#
Inside a character literal, \ introduces an escape sequence.
Supported escapes:
\n— U+000A (line feed)\r— U+000D (carriage return)\t— U+0009 (tab)\0— U+0000 (NUL)\\— backslash\'— single quote\"— double quote\xNN— a code point given as exactly two hex digits\u{...}— a code point given as 1–6 hex digits
Unicode rules:
- The decoded code point must be a Unicode scalar value:
- range
0x0000..=0x10FFFF, excluding the surrogate range0xD800..=0xDFFF. - For
\u{...}, values outside that range are rejected.
Semantics#
Evaluating a character literal produces a char value whose numeric value is
the decoded Unicode code point.
In Silk, that code point is carried as a u32 scalar.
This is an implementation detail; the language-level rule is “a char is a
Unicode scalar value”.
Examples#
ASCII and punctuation#
fn main () -> int {
let comma: char = ',';
if comma == ',' {
return 0;
}
return 1;
}
Unicode: literal UTF-8 vs \u{...}#
fn main () -> int {
let a: char = 'é';
let b: char = '\u{00E9}';
if a == b {
return 0;
}
return 1;
}
Escape sequences#
fn main () -> int {
if '\n' != '\x0A' { return 1; }
if '\r' != '\x0D' { return 2; }
if '\t' != '\x09' { return 3; }
if '\0' != '\x00' { return 4; }
if '\\' != '\u{005C}' { return 5; }
if '\'' != '\x27' { return 6; }
if '\"' != '"' { return 7; }
return 0;
}
Common Pitfalls#
- Using double quotes:
"x"is astring, not achar. Use'x'. - Writing more than one character:
'ab'is invalid; use"ab". - Source encoding surprises: prefer
\u{...}for non-ASCII characters when you want the source spelling to be stable across editors/fonts. - Confusing
\xNNbetweencharandstring: - for
char,\xNNdenotes a code point value, - for
string,\xNNdenotes a raw byte (literals string).
Related Documents#
- types (primitive
charandstring) - literals string (string literals and escape sequences)
- operators (
ascasts for int-like types, includingchar)
Source repository · Edit this page · View Markdown