Skip to content

feat: big endian c codegen - #81

Merged
hit9 merged 15 commits into
masterfrom
mpspace-io-feature/FW-1041-big-endian-c-codegen
Jun 19, 2026
Merged

hit9 merged 15 commits into
masterfrom
mpspace-io-feature/FW-1041-big-endian-c-codegen

Conversation

@hit9

@hit9 hit9 commented Jun 19, 2026 •

Copy link
Copy Markdown
Owner

Why this follow-up? The original PR (commit d117c05) only fixed opt-mode C codegen, but left three gaps:

1 Non-opt mode was still broken on big-endian hosts — the runtime (lib/c/bitproto.c) was never touched.
2 Opt-mode had a ~10% regression on little-endian — the original fix used portable shifts for everyone, punishing LE users who didn't need BE.
3 No --endian option, no tests, no docs for the new feature.

What was added

• Non-opt mode fix (lib/c/bitproto.c): Three endian-dependent sites guarded behind BP_BIG_ENDIAN. LE builds unchanged.
• Endian-dual opt-mode codegen: Emits both LE (fast byte-pointer) and BE (bit-shift) paths gated by #ifdef. Zero LE regression.
• --endian option: both|little|big to drop unused path.
• Tests: test_endian.py validating BE path on LE host.
• Makefile cleanup: Top-level clean, idempotent rm -f.
• Docs: Standalone endianness.rst + zh translations.

dcasnowdon and others added 12 commits June 16, 2026 14:52
…nter indexing

The encoder's ((unsigned char *)&(field))[fi] pattern accesses bytes via host
memory layout, which gives wrong values on big-endian targets (e.g. TMS570)
because fi=0 maps to the MSB instead of the LSB.

Replace both encoder and decoder items with explicit bit-shift arithmetic that
is endian-neutral:

  Encoder: ((unsigned_type)(chain) >> (fi*8 + shift)) & mask
  Decoder: chain |= (chain_type)(((unsigned)s[si] shift) & mask) << (fi*8))

The decoder always uses |= rather than = so it accumulates bytes without
clobbering already-written ones; a memset(m, 0, sizeof(*m)) at the top of
each generated Decode function provides the required zero baseline.
string.h is included in the generated .c file for optimization mode.

Regenerate C/CPP optimization-mode examples to reflect the new output.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Non-optimization mode relies on lib/c/bitproto.c, which copied integer
fields byte-by-byte from their in-memory representation and used
native-endian pointer-cast fast paths (`((uint32_t*)src)[0] >> si`).
Both assume a little-endian host, so encoded wire bytes were wrong on
big-endian targets.

Add a compile-time BP_BIG_ENDIAN guard (auto-detected via __BYTE_ORDER__,
user-overridable). On big-endian hosts:
- BpEndecodeBaseType stages each field through a little-endian byte view
  (reversing over the field's storage size, so e.g. uint20-in-uint32
  works), keeping the wire little-endian.
- BpCopyBufferBits skips the native multi-byte fast paths and falls back
  to the endian-neutral single-byte path.
- BpEndecodeArray disables the contiguous-memory batch copy and processes
  elements one by one.

Little-endian builds are byte-for-byte unchanged (all changes behind
#ifdef), so there is no performance impact on the common target.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The top-level example Makefile had no clean target, so stale build
artifacts (e.g. a macOS Mach-O main.o) could linger across platform
switches and break the link with "file format not recognized".

- Add a recursive `clean` to example/Makefile that cleans each C/C++
  subdir, removes the Go binaries, and sweeps lib/c/*.o.
- CPP / CPP-optimization-mode clean now also removes the shared
  lib/c/*.o object they produce.
- Make all sub-clean rules use `rm -f` so they are idempotent and don't
  error on missing files.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
assert(m1.t[i][j] = m.t[i][j]) used `=` not `==`, so the 2D-array `t`
round-trip was never actually verified (the assert checked the assigned
value, always nonzero for the test data) and it clobbered the decoded
struct. Use `==` to genuinely compare, matching the other asserts and
silencing -Wparentheses.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The portable (value-based bit-shift) opt-mode codegen is correct on
big-endian hosts but ~10% slower than the original byte-pointer code on
little-endian hosts (measured on x86-64 -O2), penalizing the common
target for a feature only big-endian hosts need.

Generate both paths guarded by #ifdef BP_BIG_ENDIAN:
- #ifndef BP_BIG_ENDIAN: the original fast byte-pointer path (decode uses
  = on first byte write, so no memset is needed).
- #else: the portable value-based path, with memset(m, 0, sizeof(*m)) for
  decode (since it always uses |=).
<string.h> is included only in the big-endian branch.

The preprocessor drops the untaken branch, so the compiled binary is
identical to single-path; only the generated .c source grows. Verified
both branches round-trip and produce byte-identical wire output, and the
little-endian compile returns to the master performance baseline.

Regenerated the C/CPP optimization-mode example artifacts.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Endianness of the host is a build-target property, not a protocol
property, so expose it as a CLI flag rather than a .bitproto option:

    --endian both    (default) emit both paths behind #ifdef BP_BIG_ENDIAN
                     and auto-detect the host via __BYTE_ORDER__
    --endian little  emit only the fast little-endian byte-pointer path
                     (smallest output, byte-identical to the pre-endian
                     codegen; no big-endian support)
    --endian big     emit only the portable big-endian value-based path

The default "both" is opt-out: unaware users get code that is correct on
any host out of the box (auto-detected, overridable via -D BP_BIG_ENDIAN),
and only someone who explicitly passes --endian=little forfeits big-endian
support in exchange for the original, smaller generated source.

Threaded through render() and BlockRenderContext; only the C renderer
consumes it. Regenerated example artifacts pick up the auto-detect snippet.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two host-independent checks, no emulation required:

- test_opt_mode_big_endian_branch: compiles each optimization-mode case
  with -DBP_BIG_ENDIAN and asserts its wire bytes match the little-endian
  build (and the Python reference when a "python" interpreter is present).
  The opt-mode big-endian path is value-based / endian-neutral, so it is
  correct even when run on a little-endian host.

- test_runtime_big_endian_layout: feeds the runtime BpEndecodeBaseType
  memory laid out by hand as big-endian and checks the little-endian wire
  bytes. Supplying the raw bytes makes it independent of the actual host
  byte order, so it exercises the host-dependent runtime path on LE CI.

A genuine big-endian end-to-end run (e.g. qemu-user s390x) remains the
gold standard and is left for CI infrastructure.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- changes.rst: add 1.3.0 entry for big-endian host support and --endian.
- performance.rst: add a "Host Byte Order (Endianness)" section explaining
  the #ifdef BP_BIG_ENDIAN dual path, auto-detection, and --endian; note the
  little-endian path is unchanged.
- compiler.rst: mention --endian in the command-line usage.
- bump __version__ to 1.3.0 to match the changelog.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@hit9 hit9 changed the title mpspace io feature/FW 1041 big endian c codegen feat: big endian c codegen Jun 19, 2026
…ions

- Move endianness content from docs/performance.rst to new docs/endianness.rst
- Add endianness to table of contents in docs/index.rst
- Update cross-references in docs/compiler.rst and docs/performance.rst
- Update zh translation files (changelog.po, compiler.po, performance.po)
- Create zh translation for the new endianness page (endianness.po)
@hit9 hit9 mentioned this pull request Jun 19, 2026
hit9 added 2 commits June 19, 2026 01:57
…mment

- docs/faq.rst: refer to big-endian support added in v1.3.0
- docs/locales/zh/LC_MESSAGES/faq.po: sync translation
- lib/c/bitproto.c: 'little-endian only' -> 'used on little-endian only'
- test_runtime_big_endian_array: BpEndecodeArray on BE processes 2 uint32
  elements independently (no batch copy), verifies encode/decode round-trip
- test_runtime_big_endian_edge_cases: uint8 (trivial), uint64 (full 8-byte
  reversal), and two consecutive uint16 fields exercising BpCopyBufferBits
  with non-zero ctx->i offset (multi-field message simulation)
@hit9
hit9 force-pushed the mpspace-io-feature/FW-1041-big-endian-c-codegen branch from a1d64df to ef80abd Compare June 19, 2026 09:13
@hit9
hit9 merged commit 0a39560 into master Jun 19, 2026
10 checks passed
@hit9
hit9 deleted the mpspace-io-feature/FW-1041-big-endian-c-codegen branch June 19, 2026 09:23
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants