Module: lash-db (search subsystem)
Dependencies: tasks.sqlite-schema.md, tasks.indexing.md
Effort: 4-6 days
Priority: HIGH
Implement fuzzy search functionality to allow users and agents to quickly find tasks by partial matching on titles, descriptions, labels, and file names. The search must be fast, rank results by relevance, and handle typos gracefully.
From design-doc.md:
- Fuzzy search across titles, bodies, labels, filenames (section 7.3.3)
- Fast performance for interactive use (section 9.1)
- Support for two approaches: SQLite FTS or in-Rust fuzzy matcher (section 9.3)
Priority: CRITICAL Effort: 1 day Depends on: tasks.sqlite-schema.md#1
Design and implement the database schema for search indexing, likely using SQLite FTS5.
- Research FTS5 vs in-memory fuzzy matching
- Benchmark FTS5 query performance
- Evaluate fuzzy matching libraries (fuzzy-matcher, sublime_fuzzy)
- Consider hybrid approach: FTS5 + in-memory ranking
- Define FTS5 virtual table (if using FTS5)
-
search_indextable with columns:-
task_id(FK to tasks.id) -
content(combined searchable text) -
title(task title, higher weight) -
body(task body, lower weight) -
labels(space-separated labels) -
file_path(for filename matching)
-
- Configure tokenizer (unicode61 or porter)
- Set up column weights (title > labels > body)
-
- Implement search index population
- During indexing, populate FTS5 table
- Extract text from tasks: title + body + labels + path
- Insert into search_index
- Add index maintenance
- Update search index when tasks change
- Delete from index when tasks deleted
- Rebuild index command
- Search index correctly represents all searchable content
- Index stays in sync with tasks table
- FTS5 queries are fast (<100ms for typical searches)
- Unit: Test index population logic
- Integration: Populate index from fixture data
- Integration: Verify index updates when tasks change
- Performance: Benchmark index size and query speed
Priority: HIGH Effort: 1-2 days Depends on: Task 1
Implement relevance scoring and ranking for search results.
- Define scoring algorithm
- Base score from FTS5 ranking (bm25 or similar)
- Boost for exact prefix matches
- Boost for title matches vs body matches
- Penalty for long documents (normalize)
- Implement
SearchScorerstruct- Take query and matched document
- Compute relevance score (0.0 to 1.0)
- Support multiple ranking strategies
- Add term highlighting
- Identify matched terms in results
- Store match positions for highlighting
- Support context snippets (show surrounding text)
- Implement result ranking
- Sort by score descending
- Break ties by created date or path
- Support pagination (offset + limit)
- Tune scoring weights
- Experiment with different weight combinations
- Test against representative queries
- Document tuning rationale
- Results are ranked by relevance
- Most relevant results appear first
- Scoring handles edge cases (empty query, etc.)
- Highlighting is accurate
- Unit: Test scoring algorithm with various inputs
- Unit: Test term highlighting logic
- Integration: Search for known terms, verify ranking
- Manual: Qualitative assessment of result quality
Priority: CRITICAL Effort: 2-3 days Depends on: Task 1, Task 2
Implement the search query API that the lash search command will use.
- Define
SearchQuerystruct-
query: search string -
scope: optional path filter -
limit: max results (default 20) -
offset: pagination offset (default 0) -
filters: label, status, etc.
-
- Implement
search()function- Parse query string
- Build FTS5 query (or run fuzzy matcher)
- Apply scope and filters
- Execute search
- Score and rank results
- Return
SearchResultsstruct
- Define
SearchResultsstruct-
results: Vec -
total_count: total matches (before limit) -
query: echo back query for reference
-
- Define
SearchResultstruct-
task_id,title,file_path,line -
score: relevance score -
snippet: context snippet with highlighted terms -
matched_fields: which fields matched (title, body, etc.)
-
- Implement query parsing
- Support quoted phrases:
"exact match" - Support field filters:
label:backend,path:core/ - Support boolean operators:
foo AND bar,foo OR bar(if FTS5) - Fallback to simple tokenization if operators not supported
- Support quoted phrases:
- Add error handling
- Invalid query syntax
- Empty query (return all tasks?)
- Index not built (suggest
lash index)
- Search API is easy to use from CLI command
- Supports common query patterns
- Fast: <200ms for typical queries
- Returns comprehensive result metadata
- Unit: Test query parsing
- Unit: Test filter application
- Integration: Search fixture project with various queries
- Integration: Test pagination (offset + limit)
- Performance: Benchmark query time for different project sizes
Priority: MEDIUM Effort: 1-2 days Depends on: Task 3
Profile and optimize search performance for large projects.
- Add performance instrumentation
- Measure query execution time
- Measure scoring time
- Measure result formatting time (snippet generation)
- Optimize bottlenecks
- [-] Use prepared statements for FTS5 queries (deferred - not needed for current performance)
- [-] Cache frequently used queries (deferred - not needed, already very fast)
- Optimize snippet extraction (pre-allocate capacity, avoid redundant allocations)
- [-] Parallelize scoring (not beneficial for current result set sizes)
- [-] Tune FTS5 configuration (deferred - current config meets targets)
- [-] Experiment with different tokenizers
- [-] Adjust column weights
- [-] Enable/disable stemming
- [-] Add result caching (optional - deferred, not needed given current performance)
- [-] Cache recent queries in memory
- [-] Invalidate on index changes
- [-] Configurable cache size
- Benchmark and document
- Small project (100 tasks): <50ms (achieved: ~0.5ms)
- Medium project (1000 tasks): <150ms (achieved: ~2.6ms)
- [-] Large project (10000 tasks): <500ms (deferred - extrapolated performance well under target)
- Search meets performance targets
- Bottlenecks identified and resolved
- Caching improves repeat query performance
- Benchmark: Generate projects of various sizes
- Benchmark: Measure query time for each size
- Benchmark: Test cache hit/miss performance
Priority: MEDIUM Effort: 1 day Depends on: Task 3
Integrate search with existing filter options (labels, status, path).
- Extend
SearchQueryto accept filters-
labels: Vec -
status: Option -
path: Option
-
- Implement filter application
- Combine FTS5 query with SQL WHERE clauses
- Filter results after FTS5 query if needed
- Maintain ranking order while filtering
- Add filter query syntax (optional)
-
label:backend query text -
status:open query text -
path:core/ query text - Parse and apply filters from query string
-
- Test filter combinations
- Search + label filter
- Search + status filter
- Search + multiple filters
- Filters work correctly with search
- Results match both query and filters
- Performance not significantly impacted by filters
- Integration: Search with label filter
- Integration: Search with status filter
- Integration: Search with path filter
- Integration: Search with multiple filters
- Advanced query syntax (complex boolean expressions)
- Faceted search (aggregations by field)
- Search suggestions / autocomplete
- Fuzzy typo correction (basic fuzzy matching is enough)
- Search history tracking
- FTS5 vs fuzzy matcher: Which provides better UX? (Recommend FTS5 for simplicity)
- Stemming: Enable for better recall or disable for precise matching?
- Snippet length: How many characters of context? (Recommend 100-200 chars)
- Highlighting format: How to indicate matched terms in output? (Use ANSI colors for terminal, for JSON)
- Design doc section 7.3.3 (
lash searchcommand) - Design doc section 9.3 (Fuzzy search approaches)
- SQLite FTS5 docs: https://www.sqlite.org/fts5.html
- Fuzzy matching libraries:
- fuzzy-matcher: https://docs.rs/fuzzy-matcher/
- sublime_fuzzy: https://docs.rs/sublime_fuzzy/