Project

token_hawk

0.0
The project is in a healthy, maintained state
Wrap LLM API calls with a single block to capture token counts, latency, and cost per call — attributed by domain and tagged with arbitrary metadata. Persists to your own Postgres database. No SaaS required.
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
 Dependencies

Runtime

>= 7.0, < 9
>= 7.0, < 9
>= 7.0, < 9
>= 1.0
 Project Readme

TokenHawk

Code-native LLM cost attribution for Rails apps.

Know exactly which feature, domain, or customer is driving your AI bill — without a SaaS dashboard, without changing how you call the API.

Gem Version CI


30-second demo

# Wrap any LLM call with a domain label
TokenHawk.track(domain: :invoice_extraction) do
  client.messages.create(model: "claude-sonnet-4-6", max_tokens: 1024, messages: [...])
end

# That's it. TokenHawk captures tokens, cost, and latency automatically.

Then from the terminal:

$ token_hawk costs --by domain

invoice_extraction    $4.21    312 calls
chat_support          $1.08     87 calls
pdf_summarizer        38¢      14 calls

Installation

Add to your Gemfile:

gem "token_hawk"

Run the generator:

bundle exec rails generate token_hawk:install
bundle exec rails db:migrate

Mount the dashboard in config/routes.rb:

mount TokenHawk::Engine, at: "/token_hawk"

Add attribution to your initializer (config/initializers/token_hawk.rb):

TokenHawk.configure do |config|
  config.storage = :active_record
end

Tracking calls

With the SDK response (automatic extraction)

TokenHawk.track(domain: :invoice_extraction) do
  client.messages.create(...)
end

Supports Anthropic and OpenAI SDK responses out of the box.

With tags (multi-tenant attribution)

TokenHawk.track(domain: :invoice_extraction, tags: { customer_id: 42 }) do
  client.messages.create(...)
end

Manual attribution (abstraction layers, PromptCanary, LangChain, etc.)

If your app wraps LLM calls behind a service layer, use record to pass token counts directly:

TokenHawk.record(
  domain:        :invoice_extraction,
  model:         "claude-sonnet-4-6",
  input_tokens:  820,
  output_tokens: 214,
  latency_ms:    1340,
  tags:          { customer_id: 42 }
)

CLI

# Monthly spend by domain (default)
token_hawk costs

# Group by day
token_hawk costs --by day

# Filter to one domain
token_hawk costs --domain invoice_extraction

# Filter to a specific tag value
token_hawk costs --tag customer_id=42

# Group by tag
token_hawk costs --by tag --tag customer_id

# Cost per call by domain (find expensive-per-call domains)
token_hawk efficiency

# Recent calls for debugging
token_hawk recent --limit 20

# JSON output (pipe to jq, etc.)
token_hawk costs --format json | jq .

Dashboard

Mount the engine and visit /token_hawk for a browser-based view:

  • Overview — monthly total, top domains by cost, daily spend trend
  • Domain detail — per-domain breakdown with top tags and recent calls
  • Recent calls — reverse-chronological call log for debugging
  • Pricing reference — supported models and their rates

The engine has no auth of its own — protect the mount point with your app's existing authentication:

# config/routes.rb
authenticate :user, ->(u) { u.admin? } do
  mount TokenHawk::Engine, at: "/token_hawk"
end

Configuration

TokenHawk.configure do |config|
  # Storage backend: :active_record (default) or :memory (tests only)
  config.storage = :active_record

  # Log telemetry failures instead of raising
  config.log_failures   = true
  config.failure_logger = ->(msg) { Rails.logger.warn(msg) }

  # Override or add pricing for unlisted models
  config.pricing["my-custom-model"] = { vendor: "anthropic", input: 0.3, output: 1.5 }
end

Subscribe hooks

React to every recorded call in real time:

TokenHawk.subscribe(:call_recorded) do |call|
  StatsD.increment("llm.calls", tags: ["domain:#{call.domain}"])
  StatsD.gauge("llm.cost_cents", call.total_cost_cents, tags: ["domain:#{call.domain}"])
end

What it doesn't do

  • No streaming support — token counts require a complete response
  • No real-time dashboard — the UI reads from the database; refresh manually
  • No multi-database support — Postgres and SQLite only in v0.1.0
  • No alerting — use subscribe hooks to wire your own

Why it exists

Most Rails apps reach for an LLM and start accumulating costs they can't explain. By the time the bill is painful, the usage is spread across dozens of call sites with no attribution. TokenHawk solves this at the source — one wrapper in your code, costs attributed from day one.


License

MIT. See LICENSE.txt.