Quirlv0.5.1
ArchitectureDecisions

Bounded automatic project discovery

Canonical Quirl project documentation synced from docs/decisions/2026-09-02_192326153_bounded-automatic-project-discovery.md.

Accepted · 2026-09-02

  • Applies to: interactive project navigation, repository discovery, and its local cache
  • Extends: ADR 0011, ADR 0016, and ADR 0021

Context

Quirl's shared picker can search files and directories, but it cannot jump directly among repositories spread across a user's machine. Requiring every user to enumerate project roots would make the common case configuration-first, while recursively scanning the home directory on the input thread would violate the shell's latency and resource contracts.

Repository discovery is also inherently stale. A repository may be created by another terminal, an editor, a version-control command, archive extraction, or a synchronized filesystem. No single command hook or filesystem notification is a complete source of truth.

The existing SQLite stores are intentionally specialized. Interactive history is durable user data, while the command-intelligence database is constructed as one complete image and atomically replaced. A rebuildable, incrementally updated project cache has a different schema, lifecycle, and failure model from both.

Decision

quirl-cli, the product composition root, owns a dedicated private projects.sqlite3 cache. Configuration supplies discovery policy and explicit roots; the database retains discovered repository paths, inferred roots, scan generations, bounded recency/frequency observations, and a cheap repository activity timestamp. Foundation, picker, process, and UI crates remain unaware of SQLite.

An interactive session starts one session-owned worker and exposes a loading snapshot while that worker admits the database and reads the bounded cache. It publishes cached repositories before starting filesystem discovery, as specified by ADR 0035. The worker performs bounded discovery independently of the input and render threads, publishes only complete SQLite transactions, wakes periodically, and accepts coalesced refresh hints. Startup, directory changes, project-picker use, successful or failed Git command completion, configuration changes, and an explicit refresh may all request work. Hints improve freshness but never replace periodic reconciliation.

Project ordering uses the newest of an explicit Quirl-open timestamp and an observed repository-activity timestamp, followed by open frequency and a stable path tie-breaker. Activity collection performs a fixed number of metadata reads for the repository directory and Git's HEAD, index, and logs/HEAD. A bounded .git file parser resolves linked-worktree metadata without launching Git. Discovery neither recursively scans repository contents nor starts one Git process per repository. Consequently, an existing file edited but not staged is represented only by Quirl visits until a future coalesced filesystem-event hint provides a stronger signal; periodic reconciliation remains authoritative for repository presence rather than every worktree-file timestamp.

Automatic discovery starts at the user's home directory plus configured roots. It checks each admitted directory for a .git directory or .git file without executing Git. A found repository is retained and its contents are pruned from the default walk. The scanner does not follow symbolic links or cross filesystem boundaries by default. It prunes platform caches, package caches, trash, build outputs, and version-control internals; explicit configuration can add roots and exclusions without silently weakening the safety defaults.

Top-level hidden directories are checked for being repositories so a repository such as ~/.dotfiles is discoverable, but arbitrary hidden trees are not recursively traversed by default. macOS ~/Library and common Linux cache and application-data trees are excluded from automatic home discovery.

Inferred roots are an explainable projection of discovered paths, not a second configuration source. The inference chooses a small non-overlapping frontier of ancestor directories that covers repository clusters while penalizing broad or expensive subtrees. Familiar names such as Code, Projects, src, and Work may prioritize the quick pass or break an otherwise equal score; names alone do not establish a root. Explicit configured roots always remain visible as such.

The rich surface exposes projects through the existing Alt-Q leader and the shared typed picker. A project selection returns its retained path to the owning CLI, which revalidates it before changing directory and preserves unfinished input. Non-interactive and degraded surfaces retain a line-oriented picker path. Project discovery is a navigation provider, not a fourth grammar mode.

Refresh and publication model

  • Cached repositories become useful before a new discovery scan completes. Worker startup or database admission failure remains an actionable diagnostic.
  • Full scans run at startup and on a bounded periodic interval. Targeted hints inspect a small set of directories and ancestors without promoting themselves to proof that an entire configured root is complete.
  • Requests use capacity-one/coalesced state. Repeated commands or timers cannot create an unbounded work queue.
  • One stable sibling lock serializes full-scan publication across Quirl processes. Interactive lock contention defers immediately; it never delays a prompt.
  • Every worker and reader owns its own SQLite connection. WAL readers may keep serving the last committed snapshot while the single writer publishes.
  • A generation may mark a repository absent only after a complete scan of that repository's owning root. Cancellation, deadline, permission failure, resource exhaustion, or lock contention preserves all previously committed rows.
  • Selection revalidates the path and repository marker. A stale row can never cause Quirl to change into a missing or no-longer-authorized directory.

Failure model and invariants

  • Filesystem input can be deep, cyclic through links, mounted remotely, mutate during traversal, contain non-UTF-8 names, deny permission, or expose an excessive number of entries. Discovery bounds depth, directories examined, repositories retained, retained path bytes, roots, total entries, and wall time. Cancellation is checked between bounded traversal turns.
  • Paths are retained in a lossless local representation. Terminal labels are separately escaped and bounded; display conversion is never used to perform a directory change.
  • A .git marker is admitted only after no-follow metadata validation. Marker contents are not trusted as commands and discovery never runs repository code, hooks, aliases, or configuration.
  • Partial observations may add a directly proven repository but may not remove another row. Only a complete owning-root generation can advance absence state.
  • Database schema, application identity, file type, permissions, and size are validated before use. Transactions and a stable lock prevent partial or stale publication; a corrupt cache is an operating error and never corrupts history or command intelligence.
  • Background work never owns terminal state, an editor buffer, a process graph, or an extension callback. Shutdown records cancellation before its bounded worker join.
  • Multiple shells may read the cache concurrently, but at most one process owns a full-scan write generation. Losing workers keep the last complete snapshot and retry only on a later bounded wake.
  • Filesystem notifications, if added later, remain lossy refresh hints. Overflow or unsupported platforms fall back to periodic bounded reconciliation.

Consequences

Project search is ready from cache at the moment users invoke it and converges without a setup wizard. First-run discovery remains local and bounded, while configured roots cover repositories outside the home directory. A separate SQLite file adds one schema and lock to maintain, but it can be rebuilt or moved aside without risking command history or the command catalog.

The traversal implementation is selected by correctness and measured cold/warm latency. The ripgrep project's ignore walker is an acceptable dependency when its pruning, cancellation, filesystem-boundary, and resource behavior remains explicit at Quirl's owning boundary; spawning an rg subprocess is not part of the runtime contract.

Performance evidence policy

The initial implementation uses one explicit std::fs breadth-first traversal rather than adding ignore. This keeps repository-root pruning, same-filesystem identity checks, symlink policy, marker errors, counters, cancellation, and the deadline in the same owning loop.

No accepted cold- or warm-scan measurements are recorded by this ADR. Results belong in the dedicated project-discovery performance record, which defines artifact identity, workload equivalence, cold/warm terminology, correctness gates, latency and cancellation measurements, and a separate ripgrep diagnostic baseline. A traversal dependency changes only after that record contains reproducible release-build evidence; unrecorded development runs are not performance claims.

On this page