Pre-Unicode and Post-Unicode Semantics

← Prev | ↑ Chapter | Next → | Index | Symbols

Pre-Unicode and Post-Unicode Semantics

Pre-2021 Autodesk Model

Before AutoCAD 2021, Autodesk documents string and character behavior in terms that were effectively tied to legacy text encodings.

In particular, Autodesk's current function notes say that earlier behavior for functions such as ascii, chr, vl-string-*, and read-char reflected ASCII or MBCS-oriented handling rather than full Unicode semantics.

Observable consequences of the pre-2021 model include:

  • character-oriented functions were effectively limited by legacy encoding assumptions,
  • file text handling depended on older platform encoding behavior,
  • string-position and character-code behavior for non-ASCII text was not Unicode-clean in the modern sense.

Post-2021 Autodesk Model

Beginning with AutoCAD 2021, Autodesk explicitly introduced Unicode-related changes controlled by LISPSYS.

From the documentation sampled here, the post-2021 model includes at least:

  • Unicode-aware string handling,
  • revised behavior for string-search and character-code functions,
  • revised read-char behavior,
  • explicit file-opening encoding support,
  • an implementation switch through LISPSYS.

LISPSYS values are documented as selecting different engine behavior:

  • legacy ASCII engine behavior,
  • Unicode engine behavior,
  • Unicode engine behavior with compatibility adjustments, depending on Autodesk release details.

For the purposes of this specification, the important normative point is:

  • string semantics after AutoCAD 2021 are versioned and mode-dependent,
  • the language cannot be specified correctly without recording the active Unicode mode.

BricsCAD

BricsCAD documentation sampled here also indicates explicit Unicode-aware text-file handling through extended open modes.

This suggests that:

  • BricsCAD should not be modeled purely as a pre-Unicode AutoLISP environment,
  • text and encoding behavior must remain part of the dialect/environment profile.

What the file-level rules actually are

The two file entry points did not gain the same thing, and the asymmetry is the source of most of the confusion in this area:

entry point encoding selected by since
open an optional third argument — AutoCAD 2021
  "utf8", "utf8-bom", "utf16-bom" BricsCAD V21
load the file's BOM alone; no argument BricsCAD V20
  exists  

In both cases, absent a BOM the host system codepage (SYSCODEPAGE) decides, and Bricsys states plainly why nothing better is possible: "UTF-8 without BOM can not be distinguished from normal ANSI."

That single sentence is the whole of the macOS mis-decoding story. A source written as UTF-8 without a mark loads correctly wherever the host codepage happens to match the author's, and silently wrong elsewhere — on macOS, where SYSCODEPAGE is MACROMAN. No error is raised at any point: the file reads, the program runs, the strings are wrong.

The portable form is UTF-8 with a BOM. See the open and load function entries for the per-function detail and the vendor sources.

clautolisp dispatch for LISPSYS

clautolisp accepts (getvar "LISPSYS") and (setvar "LISPSYS" n) under every dialect — extensions are always available, the dialect controls diagnostic emission. Per encoding-dispatch.issue:

Dialect GETVAR / SETVAR on LISPSYS
--autocad-2026 silent (native dialect)
--strict enc-extension-used
--bricscad-v26 enc-foreign-dialect / bricscad
--clautolisp enc-foreign-dialect / clautolisp

Additionally, (setvar "LISPSYS" n) with n \notin {0,1,2} emits enc-lispsys-out-of-range under all dialects; the write itself still proceeds (matching the vendor 'permissive but warn' behaviour). The diagnostics are advisory — user code that calls LISPSYS keeps running under every dialect; only the diagnosis varies. See the encoding-dispatch.issue 'Diagnostics' section for the full enc-* code table.