Repository navigation
feat: big endian c codegen - #81
Merged
Merged
Conversation
…nter indexing The encoder's ((unsigned char *)&(field))[fi] pattern accesses bytes via host memory layout, which gives wrong values on big-endian targets (e.g. TMS570) because fi=0 maps to the MSB instead of the LSB. Replace both encoder and decoder items with explicit bit-shift arithmetic that is endian-neutral: Encoder: ((unsigned_type)(chain) >> (fi*8 + shift)) & mask Decoder: chain |= (chain_type)(((unsigned)s[si] shift) & mask) << (fi*8)) The decoder always uses |= rather than = so it accumulates bytes without clobbering already-written ones; a memset(m, 0, sizeof(*m)) at the top of each generated Decode function provides the required zero baseline. string.h is included in the generated .c file for optimization mode. Regenerate C/CPP optimization-mode examples to reflect the new output. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Non-optimization mode relies on lib/c/bitproto.c, which copied integer fields byte-by-byte from their in-memory representation and used native-endian pointer-cast fast paths (`((uint32_t*)src)[0] >> si`). Both assume a little-endian host, so encoded wire bytes were wrong on big-endian targets. Add a compile-time BP_BIG_ENDIAN guard (auto-detected via __BYTE_ORDER__, user-overridable). On big-endian hosts: - BpEndecodeBaseType stages each field through a little-endian byte view (reversing over the field's storage size, so e.g. uint20-in-uint32 works), keeping the wire little-endian. - BpCopyBufferBits skips the native multi-byte fast paths and falls back to the endian-neutral single-byte path. - BpEndecodeArray disables the contiguous-memory batch copy and processes elements one by one. Little-endian builds are byte-for-byte unchanged (all changes behind #ifdef), so there is no performance impact on the common target. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The top-level example Makefile had no clean target, so stale build artifacts (e.g. a macOS Mach-O main.o) could linger across platform switches and break the link with "file format not recognized". - Add a recursive `clean` to example/Makefile that cleans each C/C++ subdir, removes the Go binaries, and sweeps lib/c/*.o. - CPP / CPP-optimization-mode clean now also removes the shared lib/c/*.o object they produce. - Make all sub-clean rules use `rm -f` so they are idempotent and don't error on missing files. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
assert(m1.t[i][j] = m.t[i][j]) used `=` not `==`, so the 2D-array `t` round-trip was never actually verified (the assert checked the assigned value, always nonzero for the test data) and it clobbered the decoded struct. Use `==` to genuinely compare, matching the other asserts and silencing -Wparentheses. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The portable (value-based bit-shift) opt-mode codegen is correct on big-endian hosts but ~10% slower than the original byte-pointer code on little-endian hosts (measured on x86-64 -O2), penalizing the common target for a feature only big-endian hosts need. Generate both paths guarded by #ifdef BP_BIG_ENDIAN: - #ifndef BP_BIG_ENDIAN: the original fast byte-pointer path (decode uses = on first byte write, so no memset is needed). - #else: the portable value-based path, with memset(m, 0, sizeof(*m)) for decode (since it always uses |=). <string.h> is included only in the big-endian branch. The preprocessor drops the untaken branch, so the compiled binary is identical to single-path; only the generated .c source grows. Verified both branches round-trip and produce byte-identical wire output, and the little-endian compile returns to the master performance baseline. Regenerated the C/CPP optimization-mode example artifacts. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Endianness of the host is a build-target property, not a protocol
property, so expose it as a CLI flag rather than a .bitproto option:
--endian both (default) emit both paths behind #ifdef BP_BIG_ENDIAN
and auto-detect the host via __BYTE_ORDER__
--endian little emit only the fast little-endian byte-pointer path
(smallest output, byte-identical to the pre-endian
codegen; no big-endian support)
--endian big emit only the portable big-endian value-based path
The default "both" is opt-out: unaware users get code that is correct on
any host out of the box (auto-detected, overridable via -D BP_BIG_ENDIAN),
and only someone who explicitly passes --endian=little forfeits big-endian
support in exchange for the original, smaller generated source.
Threaded through render() and BlockRenderContext; only the C renderer
consumes it. Regenerated example artifacts pick up the auto-detect snippet.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two host-independent checks, no emulation required: - test_opt_mode_big_endian_branch: compiles each optimization-mode case with -DBP_BIG_ENDIAN and asserts its wire bytes match the little-endian build (and the Python reference when a "python" interpreter is present). The opt-mode big-endian path is value-based / endian-neutral, so it is correct even when run on a little-endian host. - test_runtime_big_endian_layout: feeds the runtime BpEndecodeBaseType memory laid out by hand as big-endian and checks the little-endian wire bytes. Supplying the raw bytes makes it independent of the actual host byte order, so it exercises the host-dependent runtime path on LE CI. A genuine big-endian end-to-end run (e.g. qemu-user s390x) remains the gold standard and is left for CI infrastructure. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- changes.rst: add 1.3.0 entry for big-endian host support and --endian. - performance.rst: add a "Host Byte Order (Endianness)" section explaining the #ifdef BP_BIG_ENDIAN dual path, auto-detection, and --endian; note the little-endian path is unchanged. - compiler.rst: mention --endian in the command-line usage. - bump __version__ to 1.3.0 to match the changelog. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ions - Move endianness content from docs/performance.rst to new docs/endianness.rst - Add endianness to table of contents in docs/index.rst - Update cross-references in docs/compiler.rst and docs/performance.rst - Update zh translation files (changelog.po, compiler.po, performance.po) - Create zh translation for the new endianness page (endianness.po)
Merged
…mment - docs/faq.rst: refer to big-endian support added in v1.3.0 - docs/locales/zh/LC_MESSAGES/faq.po: sync translation - lib/c/bitproto.c: 'little-endian only' -> 'used on little-endian only'
- test_runtime_big_endian_array: BpEndecodeArray on BE processes 2 uint32 elements independently (no batch copy), verifies encode/decode round-trip - test_runtime_big_endian_edge_cases: uint8 (trivial), uint64 (full 8-byte reversal), and two consecutive uint16 fields exercising BpCopyBufferBits with non-zero ctx->i offset (multi-field message simulation)
hit9
force-pushed
the
mpspace-io-feature/FW-1041-big-endian-c-codegen
branch
from
June 19, 2026 09:13
a1d64df to
ef80abd
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why this follow-up? The original PR (commit d117c05) only fixed opt-mode C codegen, but left three gaps:
1 Non-opt mode was still broken on big-endian hosts — the runtime (lib/c/bitproto.c) was never touched.
2 Opt-mode had a ~10% regression on little-endian — the original fix used portable shifts for everyone, punishing LE users who didn't need BE.
3 No --endian option, no tests, no docs for the new feature.
What was added
• Non-opt mode fix (lib/c/bitproto.c): Three endian-dependent sites guarded behind BP_BIG_ENDIAN. LE builds unchanged.
• Endian-dual opt-mode codegen: Emits both LE (fast byte-pointer) and BE (bit-shift) paths gated by #ifdef. Zero LE regression.
• --endian option: both|little|big to drop unused path.
• Tests: test_endian.py validating BE path on LE host.
• Makefile cleanup: Top-level clean, idempotent rm -f.
• Docs: Standalone endianness.rst + zh translations.