brainchat
Chat with your knowledge base from the terminal, with citations that resolve to real files. brainchat is a Ruby AI CLI that answers questions over a "knowledge-brain" vault (notes, ADRs, plans, commit history, docs) and labels every claim with the exact chunk it came from.
$ brainchat ask "what did we decide about hitgate caveat timing"
Ship the two-channel honesty check first, then the self-indexed caveat [1][2]. The
rationale: publishing the caveat first would anchor the evaluation story on a limitation
with no third-party evidence behind it [1].
Sources:
[1] /vault/knowledge-brain/docs/adr/0012-caveat-investment-timing.md:1-34 (repo-docs/hitgate)
[2] /vault/knowledge-brain/memory/caveat-timing-note.md:3-9 (memory/knowledge-brain)
How it works
brainchat deliberately does not reimplement retrieval. It composes two tools that each do one thing well:
- Your question goes to
rag-query, the hybrid BM25 + cosine retrieval CLI that hitgate's regression gate benchmarks. brainchat shells out withOpen3.capture3and an argv-array (never a shell string) and treats the JSON output as the retrieval contract. - The retrieved chunks are injected into a RubyLLM chat as numbered
context. The system prompt requires
[n]citations, and the CLI prints the[n] -> path:linesmapping after the answer, so every citation resolves to a file and line range you can open.
Provider-agnostic: Anthropic, OpenAI, or fully local with Ollama.
Requirements
- Ruby >= 3.2
- The
rag-queryexecutable from hitgate on yourPATH(or pointBRAINCHAT_RAG_QUERY_BINat it), plus an indexed vault.rag-queryis configured through its ownRAG_*environment variables (e.g.RAG_INDEX_DIR), which brainchat passes straight through.-
--scope-repo allneeds hitgate at or after #134. Earlier versions turned theallsentinel into "no preference", which re-enabled cwd auto-scoping; since hitgate's source roots default to the working directory, every query got scoped to a repo named after whatever directory you were standing in and came back empty with exit 0. Ifbrainchat askfinds nothing in an index you know is populated, check this first.
-
- A provider:
ANTHROPIC_API_KEYorOPENAI_API_KEY, or a local Ollama.
Install
gem install brainchatOr build from source:
git clone https://github.com/LucasSantana-Dev/brainchat
cd brainchat && bin/setup
gem build brainchat.gemspec && gem install ./brainchat-0.1.0.gemOr run it without installing: bundle exec exe/brainchat ask "...".
Usage
brainchat ask "what did we decide about hitgate caveat timing"
brainchat ask "how does the reranker fall back" --top 8 --scope-repo all
brainchat ask "..." --provider ollama --model qwen3:4b --assume-model-exists # fully local
brainchat ask "..." --no-stream # print the answer only when complete
brainchat ask "..." --rerank # cross-encoder reranker (slower, better ranking)
brainchat version # print the gem version| Flag | Purpose |
|---|---|
--top N |
number of chunks to retrieve (default 5) |
--scope-repo a,b |
restrict retrieval to named repos, or all to disable cwd scoping |
--rerank |
enable the cross-encoder reranker (default: fast mode) |
--provider NAME |
anthropic, openai, ollama
|
--model ID |
model id (default: RubyLLM's configured default) |
--assume-model-exists |
skip model-registry validation, needed for local Ollama models |
--no-stream |
print the complete answer at once instead of streaming |
--strict-citations |
exit 1 when citation checks fail (out-of-range or missing [n] refs) |
--scores |
show retrieval scores (rrf/cos/bm25) in the Sources list |
--format json |
buffer the answer into one JSON object on stdout (answer, sources[], citation_problems[]) |
--semantic-citations |
verify cited chunks support their claims via a local judge model (experimental, warn-only) |
--judge-model ID |
judge model for --semantic-citations (default: BRAINCHAT_JUDGE_MODEL) |
Configuration
| Variable | Purpose |
|---|---|
ANTHROPIC_API_KEY |
Anthropic provider key |
OPENAI_API_KEY |
OpenAI provider key |
OLLAMA_API_BASE |
Ollama endpoint (default http://localhost:11434/v1) |
BRAINCHAT_RAG_QUERY_BIN |
path to the rag-query executable (default: rag-query on PATH) |
BRAINCHAT_RAG_QUERY_TIMEOUT |
seconds before a hung rag-query is aborted (default: 120) |
BRAINCHAT_JUDGE_MODEL |
judge model for --semantic-citations (Ollama; should differ from the chat model) |
BRAINCHAT_SEMANTIC_TIMEOUT |
seconds before the judge call is skipped (default: 120, sized for cold model loads) |
RAG_INDEX_DIR and other RAG_*
|
hitgate index configuration, passed through to rag-query
|
Under the hood
Built around Ruby idioms, not ported line-for-line from anything:
-
Open3.capture3with an argv-array, never a shell string. The question travels as one bare argv entry, so a query containing$(...), backticks or;is inert data. The spec suite proves this with a shell-injection marker test that fails if the query is ever interpolated into a shell. -
Data.definefor theChunkvalue object: immutable, keyword-constructed, with a#locationhelper that renderspath:start-end(or barepathfor line-less chunks like git commits). -
RubyLLM wired explicitly. RubyLLM 1.16 reads no provider env vars on its own, so
Chat.configurefeeds itANTHROPIC_API_KEY/OPENAI_API_KEY/OLLAMA_API_BASEitself. -
Streaming by default via
chat.ask(prompt) { |chunk| ... };--no-streamprints the complete answer at once. - Thor for CLI dispatch.
Development
bin/setup
bundle exec rake # rspec + rubocopSpecs stub both boundaries: the retriever is tested against a fixture of real rag-query
JSON output (no subprocess), and the chat against a stubbed RubyLLM (no network).
Related projects
- hitgate: the retrieval engine and regression gate brainchat queries.
- leakless: rule-driven secret scanner with redact-always AI triage, same Ruby-idiom-first approach.
License
MIT. See LICENSE.txt.