Release 0.5.44 - #541
Release 0.5.44#541
Conversation
There was a problem hiding this comment.
All reported issues were addressed across 7 files
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
Merging this PR will improve performance by 22.06%
|
| Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|
| ⚡ | int_eq/100000 |
3.2 ms | 2.6 ms | +22.06% |
Tip
Curious why performance improved? Comment @codspeedbot explain why performance improved on this PR, or directly use the CodSpeed MCP with your agent.
Comparing release/0.5.44 (0898031) with main (7005985)1
Footnotes
Benchmark Comparison (base vs PR)Measured on the same runner to eliminate hardware variance. Full resultsPerformance regressions
|
There was a problem hiding this comment.
All reported issues were addressed across 9 files (changes from recent commits).
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 18 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 32 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 23 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 14 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 11 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 14 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
1 issue found across 9 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="crates/grafeo-engine/tests/node_seek.rs">
<violation number="1" location="crates/grafeo-engine/tests/node_seek.rs:192">
P2: This wall-clock timing assertion can fail intermittently when the CI runner is loaded, and it is the only timing-based test in this file: every other test here verifies the seek deterministically via the plan (`a_key_from_the_row_is_looked_up_in_the_index` checks for `[index: id]`, `profile_shows_the_seek` checks for `NodeSeek`, and `index_persistence.rs` checks the labeled point lookup's plan). The same regression can be caught without measuring time — assert on `PROFILE` that `MATCH (s:File {id: 'n10'}) RETURN s.id` plans a `NodeSeek` and no `NodeScan`, which also removes the 60,000-node insert and the 1000-query loop from the suite. If you keep the timing check, note that both measured queries take near-identical paths in the fixed code, so the 5x ratio has a large margin; a future slowdown of the labeled path up to 5x would silently pass.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
1 issue found across 25 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="crates/grafeo-engine/src/query/planner/lpg/mod.rs">
<violation number="1" location="crates/grafeo-engine/src/query/planner/lpg/mod.rs:184">
P1: Mixed set-operation branches can bind the same output name to different value kinds, but the merged metadata records all classifications without representing per-row types. A later `RETURN` therefore applies one resolver to every branch, potentially resolving node IDs as edges; reject incompatible branch kinds or materialize each branch before combining them.</violation>
</file>
There was a problem hiding this comment.
1 issue found across 8 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="tests/spec/rosetta/values_through_ordering.gtest">
<violation number="1" location="tests/spec/rosetta/values_through_ordering.gtest:29">
P3: This ~900-character seed `INSERT` is duplicated verbatim in ten tests (all entity tests up to `a_let_copy_of_an_edge`), and the node/edge batch seed is repeated again in `skip_across_a_row_batch`/`edge_properties_after_a_cut_across_row_batches`. Every expected `w` in the file depends on this exact data, so a dataset change has to be applied in a dozen places and kept in sync with the header comment. Consider hoisting the shared seed into a `.setup` file under `tests/spec/datasets/` and referencing it per-test via the `dataset:` field. The file is otherwise correct: it parses cleanly (21 tests), all expected rows verify against the setup data, and the 2048-row batch boundary matches the engine's `DEFAULT_CHUNK_SIZE`.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
- an expand whose edge variable is already bound binds a fresh edge under an equality filter, like the check for a bound target node; CALL subqueries pass the variables they import (WITH r, WITH *) to these checks, so a cycle on imported nodes closes too - the cross join of a later MATCH keeps its inputs' column types: a node without a property no longer returns the property of the edge with the same ID, and an edge no longer reads the node's - tests: bound_edges integration tests, a subquery cycle test, a join unit test, spec bound_before_a_later_match (GQL and Cypher)
There was a problem hiding this comment.
All reported issues were addressed across 32 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 32 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 8 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 2 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
There was a problem hiding this comment.
All reported issues were addressed across 12 files (changes from recent commits).
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
What does this PR do?(tldr it adds a whole bunch of tests)
Release 0.5.44. Release checklist (the work left before the tag):
EXISTS,COUNTandVALUEsubqueries share one correlated plan inWHERE,WITHandRETURN(EXISTSandCOUNTsubqueries tied to the outer row by a value return wrong rows or fail #543), and one that shares nothing with the row runs once;CALLsubqueries return their nodes and edges as references, GQLCALLsees the outer row,CALL (a, b) { ... }scope clauses in GQL and Cypher,RETURN *inCALL, and the importingWITHfollows openCypherCALLbodies:ORDER BY,SKIP,LIMIT,UNION, a nestedCALLand writes (DELETE,MERGE,REMOVE,FOREACH) in Cypher;UNION,EXCEPT,INTERSECTandOTHERWISEin GQL; a subquery returns new names onlyCALLsubqueries andUNWIND;keys()of an edge; a variable used as both a node and an edge is rejectedOPTIONAL MATCHcondition that reads earlier variables keeps the row withnull;MERGEbinds every match; GQLNEXTpasses on what theRETURNbefore it returns; GQLVALUE { ... }inRETURNandWITHCOUNT { ... }compared inWHEREover duplicate rows, CDC events of one transaction that writes to several graphs (events now name their graph), andSETon a node deleted earlier in the same statement (now an error)compact()keeps a read-only database read-only; a failed statement inside a transaction is undone and the transaction goes onscripts/difftest/reviewed/0.5.44.txt)grafeo-cli) publishing works: a failed publish fails the release; the trusted publishers on npm and PyPI are still to be set upLater releases: writes after
compact()reaching the WAL (a known gap, noted in the CHANGELOG, #448),CREATE GRAPH ... AS COPY OF(#423), imported variables after aWITHin aCALLsubquery (#545), and a failed statement with auto-commit off and no transaction open (#536). The removal of the dormant property compression (#450) is planned for 0.5.47.Fixes #
How was it tested?
Every fix comes with a regression test that failed before it: spec tests in
tests/spec(run in Rust, Python, Node.js, Go, C# and Dart), Rust integration and unit tests, and cases in the differential test (scripts/difftest), which runs the same GQL and Cypher corpus on 0.5.43 and on this branch and lists every difference. Each change was checked withcargo fmt --check,cargo clippy --workspace --all-targets --all-features -- -D warnings,cargo test --workspace --all-featuresand the spec runners in every binding. The differential gate lists 367 differences against 0.5.43, each with its reason inscripts/difftest/reviewed/0.5.44.txt. Release-mode tests (13,803 passed), the concurrent-session stress tests and the WASM check (grafeo-web typecheck and its 345 tests against this build) pass.AI assistance
AI tools used: Claude Code for implementation, tests, review fixes and docs
Contributor checklist
Summary by cubic
Prepares the 0.5.44 release: fixes query correctness bugs (sort order, subqueries, entity typing, cyclic patterns), keeps schema/stats views correct after compaction, applies grant checks to
EXPLAINand streaming queries, and publishes npm through trusted publishing. Verified by a differential test against 0.5.43, with the CHANGELOG rewritten for the release.Bug Fixes
label_count(),property_key_count(),edge_type_count(),detailed_stats()andschema()read the graph store, and therdfprofile keeps its data across a reopen via the LPG store workaround (Persistence without thelpgfeature: RDF-only builds lose their data on reopen #544).EXPLAINand streaming queries check grants and the selected graph before running.ORDER BY,LIMIT,SKIP,DISTINCTand set operations; joins, aggregates,UNWINDandMERGEbuild output in the input columns' types.EXISTS/COUNTsubqueries inWHERErun per row when they start from a row's node and share nothing else, otherwise as a semi-join;CALLsubqueries (GQL and Cypher) return their nodes and edges as references and import the outer row;VALUE { count(x) }counts non-null distinct values; GQL subqueries chain theirMATCHclauses.MATCHafter a write sees the written rows; a path returning to an earlier variable closes the cycle; a bound edge matches only that edge; transactions read the type of edges they created even aftercompact(), concurrent direct writes aftercompact()replay to the last state, and CDC events name the graph they happened in.EXPLAINwithout parameters still shows the plan.Written for commit 0898031. Summary will update on new commits.
Closes #479
Closes #480
Closes #459
Closes #482
Closes #543