Bounded automatic project discovery
Canonical Quirl project documentation synced from docs/decisions/2026-09-02_192326153_bounded-automatic-project-discovery.md.
Accepted · 2026-09-02
- Applies to: interactive project navigation, repository discovery, and its local cache
- Extends: ADR 0011, ADR 0016, and ADR 0021
Context
Quirl's shared picker can search files and directories, but it cannot jump directly among repositories spread across a user's machine. Requiring every user to enumerate project roots would make the common case configuration-first, while recursively scanning the home directory on the input thread would violate the shell's latency and resource contracts.
Repository discovery is also inherently stale. A repository may be created by another terminal, an editor, a version-control command, archive extraction, or a synchronized filesystem. No single command hook or filesystem notification is a complete source of truth.
The existing SQLite stores are intentionally specialized. Interactive history is durable user data, while the command-intelligence database is constructed as one complete image and atomically replaced. A rebuildable, incrementally updated project cache has a different schema, lifecycle, and failure model from both.
Decision
quirl-cli, the product composition root, owns a dedicated private
projects.sqlite3 cache. Configuration supplies discovery policy and explicit
roots; the database retains discovered repository paths, inferred roots, scan
generations, bounded recency/frequency observations, and a cheap repository
activity timestamp. Foundation, picker, process, and UI crates remain unaware
of SQLite.
An interactive session starts one session-owned worker and exposes a loading snapshot while that worker admits the database and reads the bounded cache. It publishes cached repositories before starting filesystem discovery, as specified by ADR 0035. The worker performs bounded discovery independently of the input and render threads, publishes only complete SQLite transactions, wakes periodically, and accepts coalesced refresh hints. Startup, directory changes, project-picker use, successful or failed Git command completion, configuration changes, and an explicit refresh may all request work. Hints improve freshness but never replace periodic reconciliation.
Project ordering uses the newest of an explicit Quirl-open timestamp and an
observed repository-activity timestamp, followed by open frequency and a stable
path tie-breaker. Activity collection performs a fixed number of metadata reads
for the repository directory and Git's HEAD, index, and logs/HEAD. A bounded
.git file parser resolves linked-worktree metadata without launching Git.
Discovery neither recursively scans repository contents nor starts one Git
process per repository. Consequently, an existing file edited but not staged is
represented only by Quirl visits until a future coalesced filesystem-event hint
provides a stronger signal; periodic reconciliation remains authoritative for
repository presence rather than every worktree-file timestamp.
Automatic discovery starts at the user's home directory plus configured roots.
It checks each admitted directory for a .git directory or .git file without
executing Git. A found repository is retained and its contents are pruned from
the default walk. The scanner does not follow symbolic links or cross filesystem
boundaries by default. It prunes platform caches, package caches, trash, build
outputs, and version-control internals; explicit configuration can add roots and
exclusions without silently weakening the safety defaults.
Top-level hidden directories are checked for being repositories so a repository
such as ~/.dotfiles is discoverable, but arbitrary hidden trees are not
recursively traversed by default. macOS ~/Library and common Linux cache and
application-data trees are excluded from automatic home discovery.
Inferred roots are an explainable projection of discovered paths, not a second
configuration source. The inference chooses a small non-overlapping frontier of
ancestor directories that covers repository clusters while penalizing broad or
expensive subtrees. Familiar names such as Code, Projects, src, and Work
may prioritize the quick pass or break an otherwise equal score; names alone do
not establish a root. Explicit configured roots always remain visible as such.
The rich surface exposes projects through the existing Alt-Q leader and the
shared typed picker. A project selection returns its retained path to the owning
CLI, which revalidates it before changing directory and preserves unfinished
input. Non-interactive and degraded surfaces retain a line-oriented picker path.
Project discovery is a navigation provider, not a fourth grammar mode.
Refresh and publication model
- Cached repositories become useful before a new discovery scan completes. Worker startup or database admission failure remains an actionable diagnostic.
- Full scans run at startup and on a bounded periodic interval. Targeted hints inspect a small set of directories and ancestors without promoting themselves to proof that an entire configured root is complete.
- Requests use capacity-one/coalesced state. Repeated commands or timers cannot create an unbounded work queue.
- One stable sibling lock serializes full-scan publication across Quirl processes. Interactive lock contention defers immediately; it never delays a prompt.
- Every worker and reader owns its own SQLite connection. WAL readers may keep serving the last committed snapshot while the single writer publishes.
- A generation may mark a repository absent only after a complete scan of that repository's owning root. Cancellation, deadline, permission failure, resource exhaustion, or lock contention preserves all previously committed rows.
- Selection revalidates the path and repository marker. A stale row can never cause Quirl to change into a missing or no-longer-authorized directory.
Failure model and invariants
- Filesystem input can be deep, cyclic through links, mounted remotely, mutate during traversal, contain non-UTF-8 names, deny permission, or expose an excessive number of entries. Discovery bounds depth, directories examined, repositories retained, retained path bytes, roots, total entries, and wall time. Cancellation is checked between bounded traversal turns.
- Paths are retained in a lossless local representation. Terminal labels are separately escaped and bounded; display conversion is never used to perform a directory change.
- A
.gitmarker is admitted only after no-follow metadata validation. Marker contents are not trusted as commands and discovery never runs repository code, hooks, aliases, or configuration. - Partial observations may add a directly proven repository but may not remove another row. Only a complete owning-root generation can advance absence state.
- Database schema, application identity, file type, permissions, and size are validated before use. Transactions and a stable lock prevent partial or stale publication; a corrupt cache is an operating error and never corrupts history or command intelligence.
- Background work never owns terminal state, an editor buffer, a process graph, or an extension callback. Shutdown records cancellation before its bounded worker join.
- Multiple shells may read the cache concurrently, but at most one process owns a full-scan write generation. Losing workers keep the last complete snapshot and retry only on a later bounded wake.
- Filesystem notifications, if added later, remain lossy refresh hints. Overflow or unsupported platforms fall back to periodic bounded reconciliation.
Consequences
Project search is ready from cache at the moment users invoke it and converges without a setup wizard. First-run discovery remains local and bounded, while configured roots cover repositories outside the home directory. A separate SQLite file adds one schema and lock to maintain, but it can be rebuilt or moved aside without risking command history or the command catalog.
The traversal implementation is selected by correctness and measured cold/warm
latency. The ripgrep project's ignore walker is an acceptable dependency when
its pruning, cancellation, filesystem-boundary, and resource behavior remains
explicit at Quirl's owning boundary; spawning an rg subprocess is not part of
the runtime contract.
Performance evidence policy
The initial implementation uses one explicit std::fs breadth-first traversal
rather than adding ignore. This keeps repository-root pruning, same-filesystem
identity checks, symlink policy, marker errors, counters, cancellation, and the
deadline in the same owning loop.
No accepted cold- or warm-scan measurements are recorded by this ADR. Results belong in the dedicated project-discovery performance record, which defines artifact identity, workload equivalence, cold/warm terminology, correctness gates, latency and cancellation measurements, and a separate ripgrep diagnostic baseline. A traversal dependency changes only after that record contains reproducible release-build evidence; unrecorded development runs are not performance claims.
Codex CLI provides the first hosted command planner
Canonical Quirl project documentation synced from docs/decisions/2026-09-02_192326146_codex-cli-provides-the-first-hosted-command-planner.md.
Keep queued terminal input observable across resize events
Canonical Quirl project documentation synced from docs/decisions/2026-09-05_192326160_keep-queued-terminal-input-observable-across-resize-events.md.