activerecord-turbopuffer-adapter
activerecord-turbopuffer-adapter is an unofficial, fan made Ruby on Rails ActiveRecord database adapter for turbopuffer. If you are looking for the official turbopuffer Ruby gem see: https://github.com/turbopuffer/turbopuffer-ruby
The purpose of this gem is to provide Rails developers a familiar ActiveRecord style interface to turbopuffer.
Installation
To use this gem, install via Bundler by adding the following to your application's Gemfile:
gem 'activerecord-turbopuffer-adapter'Usage
Configuration
In config/database.yml define a turbopuffer adapter:
development:
adapter: turbopuffer
region: gcp-us-central1
api_key: <%= ENV["TURBOPUFFER_API_KEY"] %>
database_tasks: falseOptionally, you can define namespace_prefix, which is useful for separating namespaces for production and development environments.
Defining a model
class Document < ApplicationRecord
turbopuffer_attribute "id", "uuid", not_null: 1
turbopuffer_attribute "title", "string", filterable: true
turbopuffer_attribute "body", "string", full_text_search: true
turbopuffer_attribute "published", "bool", filterable: true
turbopuffer_attribute "embedding", "[1536]f32", ann: true
endBy default the namespace is the model's table_name, but it can be customized with self.table_name = "..." Currently, ids are always UUIDv7 (which sort chronologically, to support pagination.)
The distance metric applies to all vector columns in a namespace and defaults to cosine_distance. To use euclidean_squared instead, declare it in the model:
class Document < ApplicationRecord
turbopuffer_distance_metric "euclidean_squared"
turbopuffer_attribute "id", "uuid", not_null: 1
turbopuffer_attribute "embedding", "[1536]f32", ann: true
endAlternatively, turbopuffer can generate embeddings for you. Declare embed on a string attribute and the vector is computed on write:
class Document < ApplicationRecord
turbopuffer_attribute "id", "uuid", not_null: 1
turbopuffer_attribute "body", "string", embed: "openai/text-embedding-3-small"
endturbopuffer stores the vector in a computed embed_body attribute. embed also accepts a hash (model:, attribute:, dims:, dtype:) to name the vector attribute or set its size. Embedding is billed per token by model, see https://turbopuffer.com/docs/embedding for the supported models.
Creating records
Document.create!(title: "Hello", body: "...", published: true)
doc = Document.new(title: "Draft")
doc.save!
doc.update!(published: true)
doc.destroyNote that inserts are treated as upserts, such that writing a row whose id already exists is effectively treated as an update.
Note that transactions are not supported, interacting with that portion of the ActiveRecord API is no-op
To avoid N+1s you can use insert_all/upsert_all.
Document.insert_all(
documents.map { |doc| { title: doc.title, body: doc.body, body_embedding: doc.embedding } }
)Querying
Document.where(published: true)
Document.where(id: ["a", "b"])
Document.where.not(id: ["a", "b"])
Document.where(created_at: 1.week.ago..)
Document.where(title: /^walrus/i)
Document.where(title: Document.glob("walrus*"))
Document.where(tags: "walrus")
Document.where(tags: ["walrus", "narwhal"])
Document.where.not(tags: "walrus")
Document.where(tags: nil)
Document.where(scores: 90..)
Document.order(:title).limit(20)
Document.group(:title).count
Document.count
Document.find("018f...")
Document.rank_by("vector", "ANN", query_vector).limit(10)
Document.rank_by("text", "BM25", "quick walrus").limit(10)
Document.rank_by("body", "ANN", ["Embed", "sea mammals"]).limit(10)
Document.consistency(:eventual).where(published: true)
Document
.where(published: true)
.rank_by(["Sum", [
["Product", 2, ["category", "BM25", "mammal"]],
["text", "BM25", "quick walrus"],
]])
.limit(10)Note that turbopuffer has a limit on the maximum number of documents returned (https://turbopuffer.com/docs/query#param-limit), so Document.all.to_a etc. will only return at most 10,000 items
Eventual reads are cheaper but may not include writes from the last few seconds. Set a default with
turbopuffer_consistency "eventual"in a model orconsistency: eventualindatabase.yml. Strong is the default.
Namespaces per tenant
turbopuffer recommends one namespace per tenant: a query can only ever see that tenant's documents, and search covers just the documents that matter rather than everyone's. namespace scopes a model to any namespace, and records remember where they were loaded from, so update!, reload and destroy go back to the same place.
Document.namespace("documents-#{tenant.id}").create!(title: "Hello")
Document.namespace("documents-#{tenant.id}").where(published: true)
doc = Document.namespace("documents-#{tenant.id}").find(id)
doc.update!(published: true)The name is used as given (under namespace_prefix if one is set). Without namespace, a model's table name is its namespace.
Errors
Turbopuffer errors are raised as the ActiveRecord exceptions a Rails app already handles.
begin
Document.where(published: true).to_a
rescue ActiveRecord::ConnectionFailed
# turbopuffer is unreachable
rescue ActiveRecord::StatementTimeout
# the request timed out
rescue ActiveRecord::DatabaseConnectionError
# the API key is invalid
rescue ActiveRecord::StatementInvalid => e
e.cause # => the underlying Turbopuffer::Errors::APIError, e.g. a bad request or rate limit
endUsing alongside Postgres
While something of a lark, aspirationally the idea of this gem is to make turbopuffer conveniently usable as the primary db in a Rails app. In practice, however, using turbopuffer alongside a traditional primary db (such as postgres) is a supported, potentially more practical solution.
A dual db approach can be setup as follows:
development:
primary:
adapter: postgresql
database: myapp_development
turbopuffer:
adapter: turbopuffer
region: gcp-us-central1
api_key: <%= ENV["TURBOPUFFER_API_KEY"] %>
namespace_prefix: myapp-development
database_tasks: falseclass TurbopufferRecord < ApplicationRecord
self.abstract_class = true
connects_to database: { writing: :turbopuffer, reading: :turbopuffer }
end
class Document < TurbopufferRecord
turbopuffer_attribute "id", "uuid", not_null: 1
turbopuffer_attribute "body", "string", full_text_search: true
endContributing
Bug reports and pull requests are welcome on GitHub at https://github.com/richardmonette/activerecord-turbopuffer-adapter. As this is an unofficial gem, please do not report bugs upstream.
License
The gem is available as open source under the terms of the MIT License.