Inspiration
EveryCli started from a friction so small I ignored it for months: every time a CLI flag slipped my mind, I opened a browser tab and asked an online LLM. Each lookup cost two or three seconds of waiting and that gap was exactly long enough to lose the thread. I'd glance at a notification, switch to another task, and come back having to rebuild in my head what I was doing in the first place. The problem was never the answer quality. It was the round trip. A remote model that is smarter but three seconds away can be less useful than a small model that answers now, in the terminal I never left. So I flipped the constraint: instead of a big model reached over the network, a compact model running on my own machine, searching a local corpus I could grow myself. Less waiting, no context switch, and — critically commands I have collected and can trust.
What it does
You describe the task. EveryCli returns the command.
everycli search "how to undo my last commit" --top 2 -i
everycli search "comment annuler mon dernier commit"
- Natural-language retrieval over a curated YAML corpus, namespaced by domain — Git, Docker, Compose, npm, Composer, SSH, Python, Linux.
- Fully offline by default.
searchneeds no API key, no account, no network call. Everything is computed on-device. - Bilingual. English and French queries both work, including mixed phrasing.
- Your own command library.
everycli add,listandremovebuild a personal corpus stored separately from the shipped one, so it
survives updates and uninstalls. - Never executes behind your back. Results are shown for you to read.
--runasks for explicit confirmation. Shell wrappers only place the
command into an editable buffer — nevereval. - Built for the shell. An interactive picker (
-i),--top N,--copy,--json, and a deterministic--shellprotocol that prints only the resolved command tostdoutand sends every diagnostic tostderr. - Optional AI, clearly separated.
everycli askcalls an OpenAI-compatible API when the local corpus has no match, then offers to save the
result. Sentinel (everycli plan) is a separate Python planner that performs a safety review of a command - already - retrieved.
## How we built it
The fast path is a native Rust workspace talking to a resident daemon:
everycli-rs ──JSON over localhost TCP──▶ everycli-daemon
│ ├── YAML corpus
│ ├── model.onnx
│ ├── tokenizer.json
│ └── native ONNX Runtime
└── lexical fallback when the daemon is unavailable
everycli-core: corpus loading, YAML parsing, lexical matching, daemon discovery.everycli-inference: tokenization and embeddings through ONNX Runtime (theortcrate).everycli-daemon: a local TCP server on127.0.0.1:51821speaking one JSON line per request (ping,search,reload). It exists for one reason: loading a transformer on every keystroke is unusable, so the model stays resident in memory.everycli-rs: the CLI client, rendering, and the fallback path.
Ranking is hybrid, not purely semantic. Lexical signals catch exact tool names and flags; the semantic encoder catches paraphrase and
cross-language intent. The fused score, calibrated empirically, is:
$$ \text{score} = 0.45 \cdot \text{lexical} + 0.55 \cdot \text{semantic} + 0.20 \cdot \mathbb{1}[\text{namespace match}] $$
The namespace term is a soft route, not a hard filter — asking about Git nudges Git entries up without hiding everything else. A minimum
relevance threshold of \( 0.50 \) lets EveryCli answer "nothing relevant" instead of confidently returning noise.
The model is a fine-tuned MiniLM exported to ONNX and published at
Michelhe/everycli-minilm-ft-boosted-onnx. CI pins an exact Hugging Face
revision and verifies the SHA-256 of both model.onnx and tokenizer.json before assembling anything — the release pipeline never re-exports
the model, so a later build cannot silently ship different weights.
Distribution was treated as a first-class feature. install.sh installs into ~/.local/share/everycli and enables a systemd --user
service; install.ps1 installs into %LOCALAPPDATA%\EveryCli and configures either a Windows service or the Startup folder with -NoService.
Both verify SHA256SUMS and refuse an incomplete bundle. The end user needs no Rust, no Cargo, no Python — a moderately significant shift
from the project's original Python/PyInstaller flow.
Documentation follows a Diátaxis-style split — tutorial, how-to, reference, explanation — plus a Next.js site (App Router, Tailwind,
shadcn/ui) for the web docs and landing page.
Challenges we ran into
- The native runtime, not the Rust, was the hard part.
ort 2.0.0-rc.13requires an ONNX Runtime library from the 1.27.x series. Version 1.20.1 links but does not work with the current binary — a failure mode that costs real hours to diagnose because it doesn't look like a version problem. The workflow now pins 1.27.0 explicitly. - Platform-specific runtimes cannot be mixed.
onnxruntime.dllandlibonnxruntime.soare different files with the same job, and a bundle
that mixes them fails at load time, far from the actual mistake. Each CI job now downloads its own runtime on its own target OS. - Cold start is brutal and unavoidable. The first launch loads the model and computes embeddings for the whole corpus — minutes on a slow
machine or under WSL. The fix was honesty plus caching: the installer waits up to ~3 minutes before reporting a readiness diagnostic, and an
on-disk embedding cache makes subsequent starts fast. - Two daemons fighting over port 51821. Switching between service mode and Startup mode on Windows could leave an old instance alive. The
installer now stops the previous daemon before changing modes. - Calibrating the score honestly. Semantic-only ranking confidently returned plausible nonsense for off-topic queries. Lexical-only missed
paraphrase entirely. Getting the weights and the rejection threshold right took a real benchmark rather than intuition. - Making measurements mean anything. I kept comparing numbers that weren't comparable — a warm local fallback against a cold full daemon
round trip. The benchmarking guide now mandates stating cold vs. warm, repeating trials, and separating the first run. - macOS is not done. It compiles and passes tests in CI, but the native runtime and installer are not validated end to end, so no macOS
archive is published. Shipping an installer I hadn't verified would have been worse than shipping none.
Accomplishments that we're proud of
- A measured retrieval baseline, not a vibe. On the bilingual
eval/confusion_set.yamlbenchmark: 58 of 66 queries correct : 87.9%. - Latency where it matters. A development baseline on Windows over five trials: ~383 ms for a full client - daemon - reranked search, versus ~33 ms for the local lexical fallback.
- A genuinely zero-dependency install. Download an archive, run one script, get a working assistant with a service, a model and a runtime. No toolchain.
- Verified, not assumed. Windows install verified end to end. Linux verified under WSL across service, search and uninstall including that a normal uninstall preserves
~/.everycliand your personal commands. - A supply chain I can defend. Pinned model revision, SHA-256 verification before bundling,
SHA256SUMSpublished with each release, and
cargo auditin CI. - My own fine-tuned model, published openly rather than a black box bundled into a binary.
What we learned
- Latency is a feature, and it competes with intelligence. The right question was never "which model is best?" but "what is the largest model that fits inside the pause a developer will tolerate?" Below roughly a second, a tool feels like part of your thinking. Above three, it becomes
a tab you leave. - Hybrid beat both purebred approaches, and not marginally. Lexical matching and embeddings fail in different, complementary ways.
- A retrieval system that cannot say "no" is worse than a smaller one that can. The relevance threshold added more perceived quality than any weight tuning.
- Reproducible packaging is engineering, not chores. Pinning a revision and checking hashes prevented a whole category of "it worked yesterday" bug I would otherwise have chased blindly.
What's next for Everycli
- Finish macOS. Validate the native runtime and installer end to end, then publish the archive that CI can already build.
- Quantize the model. The current export is float32 and large. Quantization should cut download size and cold-start time the two costs users actually feel and must be evaluated against the 87.9% baseline before replacing anything.
- An ANN index for very large corpora. Exhaustive scoring is fine at today's corpus size and will not stay fine. Approximate nearest-neighbour search is the next scaling step.
- Grow the corpus with the community. The YAML schema and namespace-per-file layout were designed for contribution; the next step is making a pull request that adds twenty commands genuinely easy to review.
- Deeper shell integration, more shells, and a smoother path from "found the command" to "edited it in my prompt line."
- Sharpen the
search/askboundary, so a local miss escalates gracefully and the result flows back into your personal corpus by default every miss making the local path a little better. - Publish the web documentation site and finish the packaging polish: proper icons, deployment, and per-platform install pages.
Built With
- github-jobs
- onnx
- python
- rust
- tcp
- yaml
Log in or sign up for Devpost to join the conversation.