Blob v2 - #543
Merged
Merged
Blob v2#543
Conversation
TApplencourt
force-pushed
the
blob-v2
branch
3 times, most recently
from
September 21, 2026 22:36
c951b22 to
fac69c6
Compare
…k-around
Raw bytes -- structs, unions, opaque buffers -- were recorded as
ctf_sequence_text / ctf_array_text of uint8_t and read back with a patched
lttng-ust that emitted the full length instead of stopping at the first NUL.
lttng-ust 2.16 has first-class BLOB fields, so record them as real blobs:
ctf_array_text(uint8_t, ...) -> lttng_ust_field_fixed_length_blob
ctf_sequence_text(uint8_t, ...) -> lttng_ust_field_variable_length_blob
and map those to babeltrace blob_static / blob_dynamic. Genuine char text
(ctf_string and char sequences/arrays) is untouched: 82 fixed and 767 variable
blobs are generated, and the 10 char text fields stay as they were.
- yaml_ast_lttng.rb: the aggregate arms emit blob macros, variable when a
length_type is present and fixed otherwise.
- LTTng.rb: TracepointField learns the two blob macros, a media_type, and
blobify() for the positional dialect; print_tracepoint renders through
call_string, since a blob's arguments are not the ctf_* ones.
- gen_probe_base.rb: payload_length_field_location() names the MIP-1 field
location next to length_field_name(), which it is built from.
- gen_babeltrace_model_helper.rb: implicit_length_field? extends the
companion-length rule to variable-length blobs; blobs become
blob_static/blob_dynamic; name_packed_struct() is the one owner of "these
raw bytes are really a struct", shared by the text and blob arms.
- gen_babeltrace_lib_helper.rb: blob_static/blob_dynamic reuse the string
readback path -- a blob carries the same bytes a text sequence did.
- meta_parameters.rb: a nullable pointer-to-struct downgrades fixed to
variable so its length collapses to 0 when the pointer is null.
- babeltrace_thapi.in: request CTF 2 and MIP 1, both required to write and
read BLOB fields, when babeltrace >= 2.1.
- configure.ac: metababel >= 2.0.0, the first release that generates blob
field classes and MIP-1 field locations.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two facts about a byte buffer were spread over four places. A `void *` became a
`uint8_t` array in three copies of the same `if pointee is Void` in
meta_parameters.rb, and `yaml_ast_lttng.rb` then had to override the classifier
-- `type.name == 'uint8_t' ? :aggregate : category_of(...)` -- because uint8_t
is in INT_TYPES and so classifies as an integer. The override carried a comment
apologising for itself, and it fired 158 times: 78 arrays the generator had
synthesized and 80 the headers genuinely declare as `uint8_t *` (ze pRawData,
pKernelBinary).
Say it once on each side instead:
- TypeClasses#array_category_of(name) answers what an ARRAY of a type is.
A run of bytes is binary data -- an opaque buffer, or a struct in packed
form -- so it reports :aggregate and records as raw bytes; everything else
delegates to category_of. BYTE_TYPES names uint8_t and int8_t next to the
other fixed C type lists, and an API's own typedefs of them are found
transitively.
- MetaParameter#element_type(pointee) is the one place that says a pointer to
void points at bytes.
category_of keeps its five arms: it answers what a VALUE of a type is, and one
byte is just a number. Putting bytes there instead broke `typedef uint8_t
ze_bool_t`, which is a boolean, not a buffer -- the distinction the two
questions now draw.
int8_t is included so a signed byte buffer works if an API declares one; none
does today, and no typedef of either is used as an array element, so all 49
generated files are byte-identical.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fifty-odd declarations wrote `ctf_sequence_text, uint8_t` -- "text" for bytes
that are not text -- and LTTng.blobify rewrote each one into a blob macro on its
way past. Nothing that read those declarations could take them at face value,
and two readers did not:
- backends/ze/gen_babeltrace_ze_model.rb asked for a ctf_sequence_text row,
so gen_bt_field_model typed 175 struct-dump fields as `string` in
btx_ze_model.yaml while the tracepoints wrote blobs. The reader and the
writer disagreed about the wire.
- opencl's LTTngFieldTuple read a field's name and expression from hardcoded
tuple positions that assume a `type` slot. A blob has none, so every slot
after it shifted.
Say it once, at the source:
- the 45 YAML rows across cuda/hip/opencl/ze name the blob macro and carry a
media type; the one genuine `char` row in itt_events.yaml is untouched.
- opencl_model.rb's array broker and the two hand-written tuple sites name it
too, so blobify has no callers left and is deleted.
- LTTngFieldTuple reads slot positions from TracepointField::FIELDS, which
already declares them per macro, rather than re-encoding them.
- gen_babeltrace_cl_model.rb dispatches on the blob macro, and asks
implicit_length_field? which fields need a companion length, instead of
matching /ctf_sequence/ by name.
Also drop a dead branch: the array arm chose between a fixed and a variable
blob, but it is only reached once a length is known -- a length-less array is
logged as its address -- and the length is always a run-time size_t, so the
fixed case could not occur.
47 of 49 generated files are byte-identical. The two that move are the fixes:
btx_ze_model.yaml (175 fields string -> blob_dynamic) and opencl_model.yaml
(the macro name it records). Verified on ze_peak and a program hitting both
uint8_t shapes: structs, uuids and a 768-byte kernel binary all match the
values the program itself saw.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
These three predate the yamlfmt lint job and were never formatted. CI only lints CHANGED yaml, so editing them for the blob work is what makes them eligible -- the same way btx_tally_params.yaml surfaced. ze_events.yaml was already clean and stays clean. Pure reformatting: the parsed YAML is unchanged and all 49 generated files are byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
metababel consults this flag in exactly one place: BTFieldClass::String's setter and getter, where it means "the C variable is a struct by value, so memcpy into it rather than treating it as a char*". That is why the AST side refused to set it on a cast_type ending in `*`, under a comment admitting nobody knew why struct_names alone was not enough -- a pointer is not a struct by value, so the memcpy would have been wrong. Recording raw bytes as blobs removed the only reason it existed. All 62 fields carrying the flag are now blob_static or blob_dynamic, and the blob field classes never look at it: the static one always memcpys from `&variable`, the dynamic one from the pointer. metababel's own comment says the static blob "mirrors BTFieldClass::String with cast_type_is_struct". So THAPI was writing a flag no reader consults. That also explains a divergence: opencl set it on 11 fields whose cast_type ends in `*`, which the AST rule forbids, with no effect either way. Removing it settles the disagreement by deleting the question, and name_packed_struct no longer needs the field hash at all. All 36 generated .c files are byte-identical. Six model yamls change by exactly one thing: the key is gone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
THAPI records raw bytes as lttng-ust 2.16 BLOB fields from this branch on, so the two halves of the old work-around are gone: `55cca69.diff`, which made lttng-ust write an array of text in full rather than stopping at the first NUL, and babeltrace's null-character patch, which let the reader run past that NUL. Neither is wanted now, and the babeltrace one is actively wrong on a blob-era build -- it makes a genuine ctf_string read past its terminator. lttng-tools and babeltrace still come from the ANL branches, since those carry the pause/resume commands and the archive component class that upstream does not have; they are simply rebased onto 2.16.0 and 2.1.2. The anl-ms3-v2.1.2 branch also carries the fix for the crash behind `# TODO use anl-ms3-v2.1.2 when ctf.archive.reader doesn't segfault anymore`, so that comment goes with it, and Simon Marchi's `Use LTTNGCTL_CFLAGS`, so CI no longer downloads that patch from a pinned THAPI commit. devel pins metababel to 1.x because it emits the MIP-0 `length_field_path`. This branch is what lifts that: configure.ac requires 2.0.0, so the plain package name resolves to what it needs. babeltrace_thapi resolved `source.ctf.lttng-archive` on every invocation, which fails on a build without archive support: that component class is an ANL addition, and `~archive` links a plain babeltrace2. Resolve it only when a graph names it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
TApplencourt
force-pushed
the
blob-v2
branch
from
September 21, 2026 22:48
fac69c6 to
a8d2414
Compare
The `iprof_fs` integration test fails while reading a trace: field.rb:19:in `from_handle': unsupported field class type (RuntimeError) babeltrace2-ruby's published gem is 0.1.5, from March 2025. BLOB support was added to the project in July 2026 and has not been released, so the gem CI installs has no entry in its field-class table for the blob fields THAPI now writes, and the pretty-printer raises on the first one. Build it from the repository until a release carries those commits. That also picks up the fix for a zero-length blob, which THAPI produces whenever it traces a null pointer to a struct -- clinfo alone emits 175. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Kerilk
reviewed
Sep 22, 2026
A blob field carried "application/octet-stream" no matter what it held, so a reader could not tell one struct's bytes from another's. The 45 rows that set it said the same thing 45 times, and the media type was a positional slot, so nothing failed when a caller filled it wrongly. Derive it instead. The type a blob records is already in reach on every path: the declared events name it in the same row's `args`, the AST knows the declaration it came from, and a hand-written meta-parameter resolves `command[name]`, which its constructor already fetched and discarded. Only `blob_type=` supplies it, and one rule spells it: x86-64/zes_device_properties_t a blob of one C type application/octet-stream `void *`, which names no type The 55 rows that keep the default are the attribute queries whose length is a conditional -- their bytes really are untyped -- so the fallback is the right answer there, not a gap. Three things fall out of removing the slot. OutLTTng and InLTTng had the same constructor twice; they now share LTTngMetaParameter, the base the file's five other meta-parameter kinds already had. print_struct_tracepoint hand-spelled a whole TRACEPOINT_EVENT, media type included, and now calls LTTng.print_tracepoint. And the model helper's argument lookup moved to LTTng.argument_type, which the tracepoint side needs too. That last one was not cosmetic: the two sides had been deriving the type separately, and the model side never set it for declared events. Folding them corrects 21 media types -- 19 in ze, 2 in cuda -- that disagreed with what the tracepoint wrote. The opencl model generator never emitted media_type at all. It has the C type in hand, so it now calls the same rule. No output changes but the media type: 27 generated files differ, none by anything else. 849 blob rows in the tracepoint providers, 794 typed; 818 in the models, 763 typed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
THAPI records raw bytes as BLOB fields since the BLOB port, which bounds both halves of the round trip: lttng-ust 2.16 provides the lttng_ust_field_fixed_length_blob tracepoint macro that writes the field, and babeltrace2 2.1 the bt_field_class_blob_* API that metababel generates against to read it back. configure still asked for lttng-ust 2.10 and babeltrace2 2.0, the pre-BLOB floors. Both accept a stack that cannot build: configure passes and the build then fails on a name the headers have never heard of. Verified rather than assumed: lttng_ust_field_fixed_length_blob is absent from the 2.14 headers on this machine, and bt_field_class_blob_static_create appears first in babeltrace2 2.1.0 (absent in 2.0.6).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.