Normative Rules: Source-File and File-Stream Encoding

← Prev | ↑ Chapter | Next → | Index | Symbols

Normative Rules: Source-File and File-Stream Encoding

Statement

AutoLISP source files and host file streams have two independent encoding axes:

  1. Source-file encoding — the encoding the runtime uses when load (and any reader entry point that takes a filesystem path) reads bytes from disk. Historically AutoLISP ran on Windows with ANSI / MBCS source files, and that legacy default is observable in pre-2025 AutoCAD and pre-V25 BricsCAD. AutoCAD 2025 changed the default to UTF-8; BricsCAD V25/V26 followed suit. The value of the host system variable LISPSYS on Autodesk products selects between two engines: 0 keeps the legacy ANSI/MBCS reader, 1 and 2 enable Unicode and accept the optional encoding argument on open. LISPSYS requires an AutoCAD restart to take effect.
  2. File-stream encoding — the encoding the runtime uses when open reads or writes the bytes of a host file the user is doing I/O on. This is the third (optional) argument to open. When the argument is omitted, the per-host default applies — ANSI/MBCS for the legacy engine, UTF-8 for the modern engine.

There is no documented per-string encoding API: AutoLISP strings are sequences of code units in the host engine's native domain (8-bit code page for the legacy engine, Unicode-capable host string for the modern engine). String content is exchanged with files through the encoding configured for that file's stream.

Encoding Names

The third argument to open accepts a small documented vocabulary. Implementations should additionally accept the listed aliases since deployed source uses them interchangeably:

Designator Maps to (CL external-format)
"utf8" :utf-8
"utf8-bom" :utf-8 with byte-order mark
"utf16" :utf-16
"ANSI" :iso-8859-1 (the conservative 1-1
  mapping that never fails on bytes)
"ASCII" :ascii
"latin1" / "iso-8859-1" :iso-8859-1
"cp1252" / "windows-1252" :cp1252

Hosts derived from the BricsCAD codebase additionally accept the C-style ccs=NAME syntax embedded in the second (mode) argument, e.g. (open path "r,ccs=UTF-8"). The clautolisp bricscad-v26 dialect honours this through the existing open-ccs-mode-p knob; the strict and autocad-2026 dialects do not.

Per-Dialect Defaults

clautolisp resolves the absent third argument of open — and the absent :external-format keyword on load — using the active dialect descriptor:

Dialect Default source encoding Default file encoding
:strict :iso-8859-1 :iso-8859-1
:autocad-2026 :utf-8 :utf-8
:bricscad-v26 :utf-8 :utf-8

The strict choice is deliberately permissive at the byte level: a 1-1 byte coding never raises a decoding error on any input, and existing AutoLISP corpora produced under the legacy ANSI/MBCS engine remain loadable. Source code that needs Unicode characters in literals can either (a) use the explicit --dialect autocad-2026 / --bricscad-v26 flag, or (b) pass :external-format :utf-8 on the call site.

Vendor Evidence

  • Autodesk reference for open (AutoCAD 2024+) documents "utf8" and "utf8-bom" as the only accepted encoding strings; "When a value isn't provided for the argument, the file is assumed to contain multibyte character set (MBCS) which is the legacy behavior."
  • Autodesk LISPSYS sysvar reference: 0 = ASCII / no encoding arg, 1 = Unicode, 2 = Unicode-with-extra-COM. Restart required.
  • AutoCAD 2025 release-note quote: "In the previous versions of AutoCAD, the encoding of the file was 'ANSI', now it is 'UTF-8'." This applies to source files emitted by the Visual LISP IDE.
  • BricsCAD V26 honours the C-runtime ccs=NAME embedded mode-string as documented in earlier Phase-6 probe evidence (open-ccs-mode-p).

Per-open, not per-read

The encoding is fixed when the stream is opened, not chosen per read-char / read-line / read (or per princ / write-line). AutoCAD and BricsCAD both bind the decoder at open time — through the third argument or the ccs= mode suffix — and subsequent reads on the same handle cannot change it. clautolisp matches this: open resolves the external-format once (from its explicit argument, then *AUTOLISP-FILE-ENCODING*, then the dialect default) and holds it for the life of the stream. A program that must read parts of a file under different encodings must close and re-open it.

Extension (clautolisp): *AUTOLISP-FILE-ENCODING*

clautolisp adds a settable global, *AUTOLISP-FILE-ENCODING* — a string, where the empty string "" means "no preference; fall through to the dialect default". It is consulted when neither an explicit open third argument nor a load :external-format is supplied, and applies to BOTH directions: the read side of load / open "r" / read-line, and the write side of open "w"/"a" + princ / write-line. A pure-AutoLISP program can therefore author a file and read it back in a chosen encoding with no CLI flag. The front-end sets it from -Esource ENC (alfe) / -e ENC; it may be reassigned at runtime between open calls (not mid-stream — see the per-open rule above). This variable does not exist on AutoCAD or BricsCAD; portable code selects the encoding through the open third argument / ccs= instead. Status: clautolisp-only extension.

Measured vendor divergences (2026-07, clautolisp probe evidence)

Observed behaviours on specific (product × version × platform) targets, recorded so the divergence taxonomy (chapter 25) can classify the feature. They are not normative for the language; portable code must not rely on them. The front-end's full (situation × os × tool × version) defaults table with provenance lives in the alfe integration notes.

  • AutoCAD / accoreconsole 2022 (Windows), LISPSYS 0 — open ignores the ccs= suffix; the file-write default is ANSI1252 / MBCS, and a non-Latin-1 char is written as its low byte (U+20AC → 0xAC). LISPSYS 1/2 switch the reader/writer to UTF-8 (restart-level; LISPSYS selects cp1252 vs UTF-8). The accoreconsole console pipe is UTF-16LE, fixed by the product — no /l language flag or code page changes it.
  • BricsCAD V26 (macOS) — the open "w" default follows the POSIX locale, not SYSCODEPAGE (which reads MAC_ROMAN): a UTF-8 locale writes UTF-8, LC_ALL=C writes Latin-1. On READ, a UTF-8 file is decoded only with a BOM — a BOM-less UTF-8 file reads back as raw bytes even under ,ccs=UTF-8. ,ccs=UTF-16LE is accepted but silently writes UTF-8 — a platform divergence: on BricsCAD/Windows it writes real UTF-16LE.
  • BricsCAD V25 (Windows) — open "w" default is ANSI1252; both ,ccs=UTF-8 and ,ccs=UTF-16LE are honoured on write.
  • GUI console encodings — BricsCAD/macOS GUI console is full Unicode both ways (é = 233, € = 8364); the Windows GUI CADs (BricsCAD, AutoCAD) are cp1252 (é = 233, € = 128), AutoCAD truncating € to 0xAC on output.
  • Native (load) decode — the encoding a CAD's own (load path) uses to decode a source file follows the same (product × version × platform × LISPSYS) axis; the clautolisp front-end forwards an explicit -Esource into the CAD's per-open spelling where the engine supports it.

Vendor-documented load / open encoding rules (normative reference)

These are the concrete, version-qualified rules the vendors publish; the whole macOS cp1252 problem (a source file's (chr N) diverging from the same character read from disk) follows from them.

  • load — no encoding argument. Signature (load lspfile [failure-expr]); extension search name.des then name.lsp when none is given. Since BricsCAD V20 (and AutoCAD 2021 under LISPSYS 1/2) load also reads Unicode Lisp files stored as UTF-8 (Little-Endian) or UTF-16 (Little-Endian) — but the Byte-Order-Mark (BOM) is mandatory: a UTF-8 file without a BOM may fail. With no BOM the file is decoded with the host SYSCODEPAGE, which is read-only (e.g. MAC_ROMAN on macOS, ansi_1252 on French Windows). Consequence: a raw single-byte-codepage (e.g. cp1252) source cannot be portably native-loaded — on macOS SYSCODEPAGE = MAC_ROMAN mis-reads byte 0xE9 as È (U+00C8) instead of é (U+00E9).
  • open — optional encoding argument (AutoCAD 2021, BricsCAD V21 and higher; pre-2021 open has none). Signature (open fileName openMode [encoding]); modes "r" "w" "a" plus binary "rb" "wb" "ab". The documented encoding values are "utf8", "utf8-bom", "utf16-bom" — "UTF-8 without BOM can not be distinguished from normal ANSI". (BricsCAD additionally accepts a "MODE,ccs=ENC" mode spelling for open; ccs= is an open-only knob and has no load equivalent.)
  • Canonical-representation contract (normative) — whatever the file encoding, a decoded non-ASCII character MUST satisfy (= (chr N) <same char decoded from a file>): one canonical internal representation per character. clautolisp already guarantees this; BricsCAD's native chr is correct ((chr 233) = é) but a file char decoded with the wrong codec is not, so the two diverge — a front-end/host contract, not a chr bug. See issues/open/cad-chr-vs-loaded-encoding.issue.
  • alfe consequence — because load takes no encoding and SYSCODEPAGE is not settable, a -Esource ENC whose ENC is neither the host SYSCODEPAGE nor a BOM'd UTF-8/UTF-16 must be converted before the CAD loads it: alfe decodes the source with ENC and stages a UTF-8-with-BOM copy, which every platform's load detects and decodes correctly.

Cross-references

  • Special Form Entry: LOAD (chapter 11) — uses the dialect's source-encoding default.
  • Function Entry: OPEN (chapter 11) — uses the dialect's file-encoding default when its third argument is omitted.
  • alfe integration: the -E{situation} encoding CLI family and the (situation × os × tool × version) defaults table with provenance (source / file / console / cadstdio / log / terminal).