Repository navigation
feat(drivers): tinyflows-drivers — StateStore and Checkpointer over tinystoragedrivers ports - #108
Conversation
Add the checkpoint drivers module to the graph checkpoint package. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add the checkpoint drivers module to the graph checkpoint package. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Introduce a capability layer that exposes checkpoint drivers through the caps module, letting graphs persist and restore state via pluggable backends. The checkpoint module now routes through these drivers instead of hardcoding driver behaviour. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The driver state store now reports every storage failure as a capability error instead of splitting invalid input into a validation error, since a key the driver cannot store as an id is a capability limit rather than caller input. The storage-drivers feature also enables sha2, which the driver crate needs for its id handling. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add tests exercising the checkpoint drivers to verify state is saved and restored correctly across round-trips. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Reordered imports and reformatted long expressions in the storage driver checkpointer and its tests. The pending-write test now builds its fixture through the PendingWrite constructor instead of a hand-rolled JSON fallback, which keeps the test aligned with the type's actual API. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The driver test referenced the resume write index through the crate's public path, which fails to resolve when the test is compiled inside the crate. It now uses the crate-relative path so the test builds in both contexts. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add CHANGELOG and README entries for the new `storage-drivers` cargo feature, covering `caps::DriverStateStore` and `graph::checkpoint::DriverCheckpointer` over a tinystoragedrivers document port. The feature is off by default and requires Rust 1.88. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Tiny Sweeper reviewTiny Sweeper reviewed this change across 6 lane(s) and found 5 active actionable finding(s). Detailed lane evidence and any incomplete work are listed below. State: Ready for maintainer review Review snapshot
Completeness: Complete What changedThe review could not produce a supported behavioral summary; inspect the cited changed surface and lane details below. FeaturesNone identified with supported citations. TestsNo supported feature-to-test mapping was produced. Test execution is not inferred. Findings
Resolved this pass
Before mergeNone. Agent review detailscritique
security
tests
commits
description
e2e
Evidence and run details
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 3f3e341708
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
tinysweeper found nothing blocking. Approving.
$0.0210 · 441,459 in / 25,204 out · 42,105 cached (10%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0100 · 182,815 in / 12,358 out · 22,079 cached (12%) · gpt-5.6-luna, glm-5.3-flash
security: $0.0104 · 186,299 in / 9,288 out · 19,770 cached (11%) · gpt-5.6-luna
tests: $0.0001 · 17,049 in / 659 out · 64 cached (0%) · glm-5.3-flash
description: $0.0001 · 17,154 in / 371 out · 64 cached (0%) · glm-5.3-flash
e2e: $0.0002 · 20,562 in / 335 out · 64 cached (0%) · glm-5.3-flash
Introduce checkpoint and state store driver traits so graphs can persist and restore execution state through pluggable backends. The capabilities are wired into the driver registry and covered by tests. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The optional storage-drivers feature and its DriverStateStore re-export are gone, along with the tinystoragedrivers-core dependency they pulled in. The storage ports remain defined by the StateStore trait, so hosts can still supply their own backend without the extra crate. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Introduce a checkpoint driver and a state store abstraction so workflow execution state can be persisted and restored. The new modules ship with tests covering the checkpoint round-trip behaviour. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
The driver now overrides state_history to fetch a namespace's checkpoints and pending writes with one indexed query each, then walks the lineage in memory, replacing the trait default's two round trips per ancestor on remote backends. A by_namespace index and a stored ns_key field back the new query, and the write ledger is preferred over inline writes to match resolved_writes. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Move the merge_writes import out of the grouped tinyflows::graph use statement into its own line, matching the crate's import ordering. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
…story Add tests for deleting a thread and its pending writes, deleting named checkpoints without touching siblings or other threads, and walking a lineage newest-first in a single read including cycle handling. The history test also checks the batched result matches the trait default's hop-by-hop walk. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Replace the cloned single-element slice with std::slice::from_ref in the state history test, avoiding an unnecessary clone of the ledger write. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b4c89eb672
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
There was a problem hiding this comment.
tinysweeper found nothing blocking. Approving.
$0.0278 · 588,844 in / 40,630 out · 56,551 cached (10%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0155 · 312,767 in / 22,242 out · 33,402 cached (11%) · gpt-5.6-luna, glm-5.3-flash
security: $0.0115 · 182,977 in / 14,288 out · 18,029 cached (10%) · gpt-5.6-luna
tests: $0.0002 · 22,058 in / 573 out · 1,856 cached (8%) · glm-5.3-flash
description: $0.0002 · 22,343 in / 913 out · 1,408 cached (6%) · glm-5.3-flash
e2e: $0.0002 · 25,518 in / 913 out · 1,728 cached (7%) · glm-5.3-flash
Adds a history view over stored checkpoints and a helper for enumerating checkpoint keys, so callers can inspect past states and discover existing checkpoints without knowing their identifiers in advance. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add tests covering checkpoint save and restore behaviour in the state store, verifying that persisted state round-trips correctly. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add a state store test asserting that a document without a value field surfaces as a capability error rather than missing state, and tidy the checkpoint test module ordering and formatting. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
Add the StorageError import to the checkpoint test module so the tests can reference the storage error type. Auto-committed-on: dragonfly Co-authored-by: Medulla <medulla@tinyhumans.ai>
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7ce8740087
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| let live = self | ||
| .docs | ||
| .count(&self.checkpoints, &Filter::eq("thread", thread)) | ||
| .await |
There was a problem hiding this comment.
Avoid one remote query per retained thread counter
On a remote document store such as MongoDB, list_threads performs a serialized count request for every counter document. Because delete_thread deliberately retains these counters, the cost grows with every thread ID ever created rather than the number of live threads; a tenant with 10,000 deleted threads therefore incurs more than 10,000 round trips just to list an empty set. Derive the live thread IDs from a bulk checkpoint query or maintain their liveness without this per-counter query.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
tinysweeper found nothing blocking. Approving.
$0.0242 · 491,109 in / 40,447 out · 60,942 cached (12%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0151 · 270,626 in / 22,122 out · 36,451 cached (13%) · gpt-5.6-luna, glm-5.3-flash
security: $0.0083 · 125,000 in / 12,940 out · 19,819 cached (16%) · gpt-5.6-luna
tests: $0.0002 · 22,698 in / 989 out · 1,536 cached (7%) · glm-5.3-flash
description: $0.0002 · 22,960 in / 701 out · 1,408 cached (6%) · glm-5.3-flash
e2e: $0.0002 · 26,157 in / 932 out · 1,728 cached (7%) · glm-5.3-flash
| .await | ||
| .unwrap(); | ||
|
|
||
| store.delete_thread("doomed").await.unwrap(); |
There was a problem hiding this comment.
Exercise pending-write cleanup when deleting a thread
This test inserts checkpoints but no pending writes before deleting the thread, so it would still pass if delete_thread removed checkpoints while leaving that thread's writes orphaned. Add writes for the doomed thread and assert they are gone, while also confirming another thread's writes remain.
Additional security observation
Exercise pending-write cleanup when deleting a thread
[RULE] missing-cleanup-test
This deletion test creates checkpoints but never creates pending writes for the deleted thread. It therefore cannot catch an implementation that removes checkpoint records while leaving orphaned writes behind. Add writes for doomed, delete the thread, and assert that get_writes no longer returns them while another thread's writes remain.
[RULE] missing-delete-cleanup-coverage ·
| /// deleted; it must take that thread's writes with it and leave every other | ||
| /// thread alone. | ||
| #[tokio::test] | ||
| async fn deleting_a_thread_removes_its_checkpoints_and_leaves_others() { |
There was a problem hiding this comment.
Cover delete_checkpoints, including the empty-ids early return
The new suite exercises delete_thread but never calls delete_checkpoints. That leaves both deletion of selected checkpoint IDs and the documented empty-ID path untested, so regressions in either behavior can pass this suite. Add cases that delete selected checkpoints and verify unrelated namespaces/threads remain, including a call with an empty ID list.
Additional security observation
Cover delete_checkpoints, including the empty-ids early return
[RULE] missing-delete-test
The new tests exercise delete_thread but do not exercise the separate delete_checkpoints operation, including its empty-ID early-return path. A regression in that API could therefore pass this suite; add focused coverage for deleting selected checkpoints and for an empty ID list.
[RULE] missing-delete-checkpoints-coverage ·
| .map(|m| m.checkpoint_id) | ||
| .collect(); | ||
| assert_eq!(left, vec!["b".to_string()]); | ||
| assert!( |
There was a problem hiding this comment.
Cover delete_checkpoints, including the empty-ids early return
The empty-IDs case is now covered, but the deletion assertion checks only checkpoint a. The test deletes both a and c, so an implementation that removes c from the checkpoint list while leaving its pending writes behind would still pass. Assert that both named checkpoints have no writes after deletion.
[RULE] incomplete-test-coverage ·
| ))) | ||
| } | ||
|
|
||
| async fn get_writes(&self, config: &CheckpointConfig) -> Result<Vec<PendingWrite>> { |
There was a problem hiding this comment.
Return no writes for an unknown checkpoint
resolve_write_target returns an explicitly supplied checkpoint id without checking that the checkpoint exists, and this method then reads the write ledger directly. A caller can first call put_writes for an arbitrary id (or read after the checkpoint has been removed) and get_writes will return those writes, violating the Checkpointer::get_writes contract that unknown checkpoints return an empty vector. Verify the scoped checkpoint exists before loading its ledger.
[RULE] unknown-checkpoint-writes ·
| .iter() | ||
| .map(|tuple| tuple.checkpoint.checkpoint_id.as_str()) | ||
| .collect(); | ||
| assert_eq!(ids, vec!["c", "b", "a"]); |
There was a problem hiding this comment.
Use non-ordered IDs to test insertion ordering
The lineage IDs are inserted and expected in alphabetical order (a, b, c), so this does not distinguish walking the parent lineage from an implementation that sorts IDs in descending order. Use deliberately non-ordered IDs (for example, a root z, child a, and grandchild m) so the assertion fails for ID-based ordering.
[RULE] insufficient-test-discrimination ·
| use super::tests::{checkpoint, store}; | ||
| use super::*; |
There was a problem hiding this comment.
Start the test file with the required wildcard import
The repository rule requires test files to start with use super::*;, but the helper import precedes it. Swap these imports so the file follows the required test-file layout.
| use super::tests::{checkpoint, store}; | |
| use super::*; | |
| use super::*; | |
| use super::tests::{checkpoint, store}; |
[RULE] test-file-convention ·
| async fn deleting_checkpoints_removes_only_the_named_ones_and_their_writes() { | ||
| let saver = store(); | ||
| assert_eq!(saver.delete_checkpoints("t", &[]).await.unwrap(), 0); | ||
| for (thread, id) in [("t", "a"), ("t", "b"), ("t", "c"), ("other", "a")] { |
There was a problem hiding this comment.
Use non-ordered IDs to test insertion ordering
The deletion fixture inserts checkpoint IDs in lexical order and only asserts that the surviving ID is b. This cannot expose an implementation that accidentally derives ordering from the identifier or insertion sequence. Use deliberately non-ordered IDs and assert the documented ordering explicitly.
[RULE] insufficient-ordering-coverage ·
| mod checkpoint_keys; | ||
| mod state_store; | ||
|
|
||
| pub use checkpoint::{DEFAULT_PREFIX, DriverCheckpointer}; |
There was a problem hiding this comment.
Drive a flow through the driver checkpointer and state store end to end
DriverStateStore and DriverCheckpointer are host-facing adapters whose whole point is to be plugged into engine::run_with_checkpointer / resume_with_checkpointer and Capabilities::state. No end-to-end test reaches that: the harness is entirely in crates/tinyflows and never constructs a DriverCheckpointer or DriverStateStore (the candidate hits are lexical matches on words like source, thread, namespace), and there is no e2e workflow in the tree at all. Every guarantee the new code makes — CAS seq reservation, namespace-scoped resume, write-ledger precedence in state_history — is currently verified only against the adapter's own trait methods, not against a flow that runs, checkpoints, and resumes through it. A test would have to run a small graph with the driver checkpointer and state store wired in, complete a superstep, then resume on a fresh handle from the same backing store and observe the flow continue from the persisted checkpoint/state.
[RULE] e2e-uncovered ·
Co-authored-by: Medulla <medulla@tinyhumans.ai>
There was a problem hiding this comment.
tinysweeper found nothing blocking. Approving.
$0.0013 · 102,041 in / 5,603 out · 2,918 cached (3%) · gpt-5.6-luna, glm-5.3-flash
critique: $0.0004 · 6,740 in / 1,285 out · 2,598 cached (39%) · gpt-5.6-luna, glm-5.3-flash
tests: $0.0002 · 22,741 in / 1,259 out · 64 cached (0%) · glm-5.3-flash
description: $0.0002 · 23,003 in / 926 out · 64 cached (0%) · glm-5.3-flash
e2e: $0.0002 · 26,200 in / 800 out · 64 cached (0%) · glm-5.3-flash
| `state_history` in two queries, length-prefixes every key component and | ||
| hashes over-long ids with SHA-256. A separate crate rather than an engine | ||
| feature because the published `tinyflows` package cannot carry a git | ||
| dependency; the engine is unchanged and keeps its MSRV, this crate needs |
There was a problem hiding this comment.
Keep the new crate compatible with the repository MSRV
The repository rules require Rust 1.85, but this entry explicitly introduces a workspace crate requiring Rust 1.88. That makes the new crate incompatible with the declared MSRV for users and CI environments pinned to 1.85. Either lower the crate's requirement to Rust 1.85 or update the repository's MSRV policy and all affected compatibility guarantees together.
[RULE] msrv-compatibility ·
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 9fd522e1a2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| Some(found) => ( | ||
| found.doc.get("next").and_then(Value::as_u64).unwrap_or(0), | ||
| found.unchanged(), |
There was a problem hiding this comment.
Reject malformed counters instead of resetting the sequence
When an existing thread-counter document has a missing, negative, or non-integer next, this silently treats it as zero and successfully overwrites the counter before inserting the checkpoint. For a thread with existing checkpoints, subsequent puts then collide with old low sequence IDs or create a checkpoint ordered behind the actual latest record, so a successful write may be invisible to resume. Return a checkpoint error for a malformed existing counter rather than mutating it as though the thread were new.
Useful? React with 👍 / 👎.
Summary
This is the first tinyflows step of the storage-drivers migration. A new unpublished crate,
tinyflows-drivers, adds two backends for the engine over a tinystoragedrivers document port. The publishedtinyflowsengine is unchanged (no diff incrates/tinyflows). A host can then keep flow-run state and graph checkpoints in whichever backend it opened (SQLite on a desktop, MongoDB in the cloud, memory in tests), scoped per tenant.DriverStateStore, aStateStore. Each key is one document{key, value}inflows_state. Keys the driver cannot use as ids are refused as capability errors rather than hashed, so two keys never share a slot.DriverCheckpointer<State>, a graphCheckpointer. It is ported from TinyAgents' mergedDriverCheckpointer(feat(storage): tinystoragedrivers adapters for the harness store and graph checkpointer tinyagents#333), including everything its two review rounds hardened:seq, so listing is insertion order and the latest write wins for a reused id, as in the SQLite and file backends;get_scopedpushed down to one indexed(thread, namespace, seq)query;merge_writesunder compare-and-swap;state_historythat reads the namespace's checkpoints and pending writes in two queries and walks the lineage in memory, instead of two round trips per ancestor.The lease methods are not ported, because tinyflows' trait has none.
Default collections (
flows_state,flows_graph_*) stay clear of TinyAgents'graph_*collections, so both can share one scoped backend.Dependency:
tinystoragedrivers-coreis a git dependency at tagv0.4.0, not a vendored copy. A host that already vendors the storage crates (OpenHuman, through tinyagents) points that URL at its one copy with[patch."https://github.com/tinyhumansai/tinystoragedrivers"], the same way it unifiestinytools. Then there is exactly oneDocumentStoretrait in the graph.Why a separate crate: the published
tinyflowspackage cannot carry a git dependency, even an optional one, socargo publishwould fail.tinyflows-driversispublish = false, needs Rust 1.88 (the storage crates' MSRV), and leaves the engine's dependencies and its 1.85 MSRV untouched.cargo package -p tinyflowsstill succeeds.Next steps: put
tinyflows-sqlite'sflows.db/jobs.dbon the SQLite driver's native mode, and shimtinyflows-adaptive'sStorageonto the storage config and scopes.Tests
DriverCheckpointerpasses the same behavioural checks as the SQLite checkpointer, ported fromtinyflows-sqlite'scheckpoint_tests.rs:Added on top:
Bulk
state_history: newest-first lineage, off-lineage siblings and other namespaces excluded, the write ledger preferred,limit, a cycle guard, and agreement with the trait default's hop-by-hop walk.delete_threaddrops pending writes;delete_checkpointsremoves only the named ids and their writes, leaving other threads alone (plus its empty-ids early return).DriverStateStore: round trips, a storednull, scope and collection isolation, and unstorable keys.