SQA::BI
Bayesian inference over discrete outcomes from time series data.
SQA::BI answers one question: given this feature vector, what is the
probability distribution over outcomes? โ where "outcomes" is a small
ordered set such as [-2, -1, 0, 1, 2], read as strong downtrend through
strong uptrend.
It gives you a full posterior rather than a point estimate, so entropy, confidence, and information gain come out of the math instead of being bolted on afterward.
๐ Full documentation โ guide, API reference, LLM integration, and runnable examples.
โ ๏ธ WARNING: This is a learning tool, not production software. DO NOT use this library when real money is at stake. The probability distributions it produces should not be taken seriously. If you lose your shirt playing in the stock market, don't come crying to me. Playing in the market is like playing in the street โ you're going to get run over.
Where it sits in the workspace
sqa-tai โโ
sqa-bi โโดโ sqa โ sqa-cli / sqa-advisor / sqa-rails / sqa-sinatra
sqa-bi is a leaf gem with no runtime dependencies. The inference math
is domain-agnostic โ it knows nothing about markets โ so sqa can depend on
it the same way it depends on sqa-tai. The market-facing demo
(examples/03_stock_market_prediction_v2.rb) reaches the other direction
and needs sqa, which is why Gemfile.local resolves sqa from the
sibling checkout for development only.
Quick start
require "sqa/bi"
predictor = SQA::BI.predictor(outcomes: [-2, -1, 0, 1, 2], bandwidth: 1.0)
predictor.train([1.0, 2.0, 3.0], outcome: 1)
predictor.train([1.1, 2.1, 2.9], outcome: 1)
predictor.train([-1.0, -2.0, 0.5], outcome: -2)
posterior = predictor.predict([1.05, 2.05, 3.0])
posterior.max_outcome # => 1 (MAP estimate)
posterior.probability(1)# => 0.74
posterior.confidence # => 0.41 (1 - entropy / max entropy)
posterior.entropy # => 1.37 bits
puts posterior.summaryComponents
| Class | Role |
|---|---|
Prior |
P(outcome) โ uniform or custom, Laplace-smoothed updates from observed frequencies, weighted combination, entropy |
Likelihood |
P(features | outcome) โ Gaussian kernel density estimation over historical observations |
Posterior |
P(outcome | features) โ Bayes' theorem, plus entropy, confidence, KL divergence from the prior, MAP, sampling |
TimeSeriesPredictor |
the train / predict interface tying the three together |
LlmPriorElicitor |
an LLM supplies the prior from a natural-language description |
LlmLikelihoodEstimator |
an LLM acts as a likelihood function over textual evidence |
LlmSupport |
provider resolution, JSON extraction, re-keying, normalization, clamping |
Bayes' theorem: P(outcome | data) = P(data | outcome) ร P(outcome) / P(data)
Kernel density estimate for the likelihood:
P(x | outcome) = (1/n) ร ฮฃ K((x - xแตข) / h), with K Gaussian and h the
bandwidth. Smaller bandwidth tracks local patterns; larger smooths.
LLM integration
The design rule, stated once and applied everywhere:
The LLM judges. Ruby computes. Ask the LLM only for isolated, independent judgments (a weight, a conditional probability). Never ask it to accumulate, normalize, or update beliefs โ that is Bayes' job, done in Ruby.
LLMs and Bayes' theorem have exactly complementary failure modes: an LLM is
good at "how surprising is this log line if the database were down?" and bad
at combining five such judgments without anchoring or double-counting. The
posterior is good at exactly the second thing and has no opinion about the
first. See docs/EXPLORATION.md for the verified
results behind both patterns.
Prior elicitation answers the standing objection to Bayesian methods โ where does the prior come from? โ by treating the LLM as a queryable compression of domain knowledge:
elicitor = SQA::BI::LlmPriorElicitor.new(
outcomes: [-2, -1, 0, 1, 2],
outcome_descriptions: {
-2 => "strong downtrend", -1 => "mild downtrend", 0 => "sideways",
1 => "mild uptrend", 2 => "strong uptrend"
}
)
prior = elicitor.elicit("The Fed unexpectedly cut rates by 50bp...")Likelihood estimation handles evidence that is text rather than numbers:
estimator = SQA::BI::LlmLikelihoodEstimator.new(
hypotheses: {
bad_deploy: "The 14:02 deploy introduced a bug",
database: "The primary database is degraded",
network: "There is a network partition between AZs"
}
)
estimator.likelihoods("Error rate spiked 2 minutes after deploy")
# => { bad_deploy: 0.9, database: 0.2, network: 0.15 }Likelihoods are clamped to [0.001, 0.999] so an overconfident model can
never zero out a hypothesis in a single step (Cromwell's rule) โ which is
what keeps later contradicting evidence able to reverse the belief.
ruby_llm is required lazily inside LlmSupport.build_chat, so the core
math loads and runs without it. Both LLM classes accept an injectable
chat: object; the test suite makes zero network calls.
Provider selection โ local first
Auto-detected in order: LM Studio via ruby_llm-providers-lms
(http://localhost:1234/v1), then Apfel / Apple Foundation Models via
ruby_llm-providers-apfel (http://127.0.0.1:11434/v1), then cloud.
Detection probes each server, so a local provider is only used when its server is actually running with a chat model available. For LM Studio that means two things, neither of which the desktop app does by itself:
lms status # Server: ON / OFF
lms server start # start the HTTP server on :1234
lms ps # which models are loaded
lms load qwen/qwen3.8-27b # load one (choose_local_model prefers qwen)With the server off, or on but offering only embedding models, resolution falls through to cloud โ silently, by design. That is how an expected local run turns into an unexpected cloud bill or a 401 from a stale key. The LLM demos therefore print where they are actually sending each question before the first call:
LLM: lms โ qwen/qwen3.8-27b at http://localhost:1234/v1
SQA::BI::LlmSupport.current_resolution returns that line if you want the
same check in your own code.
| Variable | Effect |
|---|---|
SQA_BI_LLM_PROVIDER |
force lms, apfel, or cloud
|
SQA_BI_LLM_MODEL |
force a model id |
LMS_API_BASE, APFEL_API_BASE
|
override the local server URLs |
The unprefixed BI_LLM_PROVIDER / BI_LLM_MODEL names this library used
before it moved into the SQA workspace are still honored as a fallback.
For local servers, choose_local_model prefers qwen models โ they honor
JSON prompts most reliably in LM Studio โ then gpt-oss, skipping embedding
and OCR models. Model quality shows up directly as prior sharpness: qwen3.8-27b
via LM Studio produced a bullish prior peaked at +1, while Apple's 3B
on-device model was more cautious and peaked at 0.
Examples
ruby examples/01_coin_flip.rb # the mechanics, minimal
ruby examples/02_time_series_prediction.rb # synthetic trends
ruby examples/03_stock_market_prediction_v2.rb # real OHLCV via sqa
ruby examples/04_llm_elicited_prior.rb # LLM as prior elicitor
ruby examples/05_llm_likelihood_diagnosis.rb # LLM as likelihood functionExample 03 needs the sqa gem, so run it in dev bundle mode
(asgard dev from the workspace root, which points BUNDLE_GEMFILE at
Gemfile.local). examples/pure_ruby_indicators.rb carries dependency-free
SMA/EMA/RSI so the other examples stay standalone โ production code should
use sqa-tai instead.
Documentation
The detailed docs are an MkDocs site under docs/:
pip install -r docs/requirements.txt
mkdocs serve # live reload at http://127.0.0.1:8000
mkdocs build --strict--strict turns the link and anchor validation configured in mkdocs.yml into
build errors, so a broken cross-reference fails the build rather than shipping.
.github/workflows/docs.yml publishes to GitHub Pages on push to main.
Diagrams are hand-authored SVG in docs/assets/diagrams/ โ dark theme,
transparent background, colour used to distinguish function.
Development
asgard test # or: bundle exec rake test
asgard quality # tests + coverage, RuboCop, Flog, Flay, Reek
asgard rubocopReek runs baseline-aware against .quality/reek_baseline.txt; ratchet the
floor down with asgard reek_baseline after a genuine improvement. The
RuboCop config is generated from the workspace's .rubocop.yml.common โ
edit that file and run asgard sync_rubocop, not this repo's copy.
History
Ported from the bayesian_inference prototype in
~/sandbox/git_repos/madbomber/experiments/ai_misc/. See
decision_support_techniques.md for the
investigation log that led here, including why a tree-ensemble forecaster
and the Laya typed-decision model were each evaluated and set aside.
License
MIT. See LICENSE.txt.