# struct-reader **Repository Path**: lp9906/struct-reader ## Basic Information - **Project Name**: struct-reader - **Description**: 把一个文件完整映射到一个普通 class的结构化读取器 - **Primary Language**: Unknown - **License**: MIT - **Default Branch**: main - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-09-29 - **Last Updated**: 2026-09-29 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # StructReader 把**一个文件完整映射到一个普通 class(非静态)**的结构化读取器。 用链式声明字节布局,用 `DataView` + 游标(`StructStream`)顺序读取, 读出的数据落在你自己定义的 class 的 **public 实例属性**上。零运行时依赖,`ES2022+`。 ## 安装 ```bash pnpm add struct-reader ``` ## 快速上手 ```ts class Sprite { public magic = ''; public frameCount = 0; public frames: number[] = []; } const sprite = new StructReader(Sprite) .field('magic', atoms.chars(4)) .field('frameCount', atoms.u16) .array('frames', ctx => ctx.current.frameCount, atoms.u16); const data = sprite.read(buffer); data.magic; // 'SPRT' data.frames; // [0x10, 0x20] ``` ## 两类入口 | | 类 | 用途 | | --- | --- | --- | | **本层** | `StructReader` | 每个方法声明一个**具名字段** | | **字段里** | `StructChain` | 已经站在某个字段上,只声明它的形态 | **`StructReader`** | 方法 | 作用 | | --- | --- | | `.field(name, read)` | 读一个标量字段 | | `.field(assign, read)` | 读出后用回调自己落位 | | `.struct(name, Ctor, declaration)` | 读一段结构,实例由构造器产出 | | `.switch(name, branch => …)` | 按条件选一条分支 | | `.when(name, condition, declaration)` | 条件成立才读这一段 | | `.ptr(name, getOffset, read \| declaration)` | 跳到偏移处读一个值,或继续声明 | | `.array(name, getLength, read \| subCall)` | 读一段数组 | | `.list(name, getSize, read \| subCall)` | 读一段列表,`getSize` 给窗口字节数 | | `.repeat(name, read \| subCall)` | 重复读元素直到窗口结束 | | `.range(getSize, declaration)` | 在限定窗口内继续声明本层 | | `.lazy(subCall)` | 这段声明延后到需要时再读 | **`StructChain`** —— 同样的形态,省掉名字:`struct` / `switch` / `ptr` / `array` / `list` / `repeat` / `range` / `lazy`。 ## 读法与落点 **读法(`Atom`)是带标记的函数**,只能由 `atom()` 造出来: ```ts export type Atom = ((stream: StructStream) => Value) & { readonly _atom: true }; export function atom(read: (stream: StructStream) => Value): Atom { return Object.assign(read, { _atom: true as const }); } ``` 标记的用处是**分辨同位置的两种形态**:`array(name, len, read)` 与 `array(name, len, subCall)` 都是「一个函数」,只有标记分得开。`isAtom()` 在运行时做这个判断,所以各处入口能按有无标记分流。 词表外的读法自己包一层即可: ```ts .field('flags', atom((stream) => stream.u16() & 0xff)) ``` **落点默认是字段名**,也可以给一个回调自己决定往哪写: ```ts .field((current, value: number) => { current.first = value; current.second = value * 2; }, atoms.u8) .array((current, value: number[]) => current.all.push(...value), 2, atoms.u8) ``` 回调收到的 `current` 是本层产物、`value` 是读法产出,类型由读法推出来。 ## 尺寸 `getLength` / `getSize` 可以给函数,也可以直接给数字: ```ts .array('fixed', 2, atoms.u8) // 常量 .list('chunk', ctx => ctx.current.size, atoms.u8) // 按已读字段 ``` ## 读取与挂起 ```ts const data = reader.read(buffer, env); // 返回产物本身 ``` `buffer` 是 `ArrayBuffer | ArrayBufferView | DataView`;`env` 在 `TEnv` 为 `void`(默认)时可省略。 `.lazy` 是顺序断点,`read()` 到此为止,余下交给 `reader.pending`: ```ts const reader = new StructReader(Chunk) .field('head', atoms.u8) .field('tail', atoms.u8) .lazy(inner => inner.field('middle', atoms.u8)) .lazy(inner => inner.field('extra', atoms.u8)); const data = reader.read(buffer); data.middle; // 0 —— 还没读 reader.pending.hasNext(); // true reader.pending.getPending(); // ['middle', 'extra'] reader.pending.next(); // 读第一段 data.middle; // 3 reader.pending.next(); // 读第二段 reader.pending.hasNext(); // false ``` `next()` 按声明顺序从当前流位置接着读;`next(path)` 按路径精确取指定的一段 —— 元素内的 `lazy` 包 `ptr` 时,可以跳过前面的元素直接解某一个: ```ts catalog.pending.next('entries[1].body'); ``` ## atoms | 原子 | 说明 | | --- | --- | | `u8` `i8` | 1 字节整数 | | `u16` `i16` `u16be` `i16be` | 2 字节,默认小端,`be` 后缀为大端 | | `u24` `i24` `u24be` `i24be` | 3 字节 | | `u32` `i32` `u32be` `i32be` | 4 字节 | | `u64` `i64` `u64be` `i64be` | 8 字节,读出 `bigint` | | `f32` `f64` `f32be` `f64be` | IEEE754 浮点 | | `bool(count)` | 读 `count` 位,非 0 为真 | | `bits(n)` | 读 n 位,低位在前,跨字节续读 | | `chars(n)` | 定长字符,原样返回 | | `str(encode, n?)` | 定长字符串,去掉尾部 `\0`;省略 `n` 读到 `\0` 为止 | | `str_utf8(n?)` `str_latin1(n?)` | 同上,指定编码 | | `bytes(n)` | 零拷贝 `Uint8Array` | | `skip(n)` | 跳过 n 字节 | | `flag(FlagConstructor, read)` | 读整数并包成 `Flag` 子类实例 | | `enum(members, read)` | 读一个值,是枚举成员就返回它,否则 `undefined` | 原子都是纯读法,**只看流**。需要动态长度或已读字段时,写在收 `context` 的位置(取长 / 取向 / 条件回调)。 ## 回调上下文 ```ts class StructContext { public stream: StructStream; public current: TCur; // 本层已读的字段 public parent: TParent | null; // 父层的上下文 public root: TRoot | null; // 顶层产物,读完才回填 public env: TEnv; // read(buffer, env) 传进来的环境 public path: string; // 'frames[2]' } ``` 四类回调会收到它:取长(`getLength` / `getSize`)、取向(`getOffset`)、条件(`when` 与 `switch` 的条件)。 **读法不收 context** —— 所以 `env` 的作用是选择走哪条分支,不是改变怎么读: ```ts const scaled = new StructReader(Scaled) .when('value', ctx => ctx.env.scale === 10, inner => inner .field('value', atom((stream) => stream.u8() * 10)) ) .when('value', ctx => ctx.env.scale !== 10, inner => inner .field('value', atom((stream) => stream.u16())) ); ``` ## 例子 **结构字段:实例由构造器产出** ```ts const framed = new StructReader(Frame) .struct('rect', Rect, item => item .field('x', atoms.i16) .field('y', atoms.i16) ) .field('durationMs', atoms.u16); const data = framed.read(buffer); data.rect; // Rect 实例 data.rect.x; // -1 ``` **条件分支** ```ts const variant = new StructReader(Variant) .field('version', atoms.u8) .switch('body', branch => branch .case(ctx => ctx.current.version >= 2, atoms.u32) .default(atoms.u16) ) .when('extra', ctx => ctx.current.flag === 1, inner => inner .field('extra', atoms.u8) ); ``` `switch` 取第一个条件成立的分支;分支可以直接给读法,也可以给声明回调(写多个字段)。 **数组元素:读法或声明** ```ts const slotted = new StructReader(Slotted) .field('count', atoms.u8) .array('slots', ({ current }) => current.count, slot => slot .field('id', atoms.u8) .field('nameOffset', atoms.u8) .ptr('name', ctx => Number(ctx.current.nameOffset), atoms.str_utf8()) ); ``` 元素是标量时第三位直接给读法(`.array('frames', len, atoms.u16)`),要声明多个字段时给回调。 **目录表:元素内的 `lazy` 包 `ptr`** ```ts const catalog = new StructReader(Catalog) .field('magic', atoms.chars(4)) .field('count', atoms.u32) .array('entries', ({ current }) => current.count, entry => entry .field('name', atoms.str('utf-8', 8)) .field('offset', atoms.u32) .lazy(inner => inner .ptr('body', ctx => Number(ctx.current.offset), atoms.str_utf8()) ) ); const data = catalog.read(buffer); data.entries[0].name; // 'a.txt' data.entries[0].body; // undefined 该元素的 lazy 段还没跑 catalog.pending.getPending(); // ['entries[0].body', 'entries[1].body'] catalog.pending.next('entries[1].body'); // 跳过第一个元素,直接解第二个 data.entries[0].body; // undefined 仍挂起 data.entries[1].body; // 'world' ``` **位打包** ```ts .lazy(inner => inner.array('tiles', ({ current }) => current.mx * current.my, tile => tile .field('groundHeight', atoms.i16) .field('waterLevel', atoms.bits(14)) .field('waterBoundary', atoms.bool(2)) .switch('groundTextureIndex', branch => branch .case(ctx => ctx.root?.version === W3eVersion.REFORGED, atoms.bits(6)) .default(atoms.bits(4)) ) .field('flags', atoms.flag(PackedFlag, atoms.bits(10))) )) ``` **标志位、枚举、角度** ```ts .field('flags', atoms.flag(SpriteFlags, atoms.u16)) .field('kind', atoms.enum(ChunkKind, atoms.str('ascii', 4))) .field('spin', radians(atoms.f32)) ``` ## 越界与路径 越界抛 `StructRangeError`,错误信息带字段路径: ``` frames[1] out of range at offset 0x8 (need 2, available 1) ``` 路径由读取过程拼出:进字段加 `.name`,进元素加 `[index]`。用自定义落点(`assign` 回调)的字段没有名字,报错就只显示所在层。 ## License [MIT](LICENSE)