Project

agentilda

0.0
The project is in a healthy, maintained state
Keeps a project's specifications, plans and pull requests joined up, and drives specialist agents over them in parallel. A feature's state is its folder name under .plans, so a transition renames a directory rather than updating a row, and nothing can claim a phase whose document is missing.
2005
2006
2007
2008
2009
2010
2011
2012
2013
2014
2015
2016
2017
2018
2019
2020
2021
2022
2023
2024
2025
2026
 Project Readme

Agentilda

CircleCI Coverage

What is this?

This repo is a collection of installers (BASH and Ruby) which ensure a consistent vendor-neutral agentic setup on your computer, coupled with an agentic software team that diligently creates your feature specs in collaboration with you, and then works on them until they become reviewed, CI-passing PRs that you get to merge.

Here is a screenshot of agents working on two plans at the same time, but on two different phases of the process:

workflow

Note

  1. agentilda (also known as tilda executable) is the Ruby Gem, which is a CLI tool that creates and manages the .plans folder, and comes with eight or so specialized agents that take a spec.md file and work through it until it's a set of PRs open, reviewed, and passing on your CI. It does not automatically merge anything.
  2. agentilda-ai-setup is the GitHub repo that's a mixture of BASH and Ruby installers. It's comes with the configuration.yml file, which lists the installation commands for the coding agents you'd like to install locally, any other executables you might want (for instance, it installs bt — braintrust's CLI utility), and then you can list any number of Github Repos and use them to install skills, plugins, commands from them, specifying exactly which you want to install and which you want to exclude. Moreover you can specify a sub-directory of a github repo to install from.
  3. The final piece of the puzzle is the locking gem agent-lock. This flexible gem comes with the executable alock and a skill teaching agents how to use it. Using alock agents can work in parallel in the same worktree but on different files, sub-folders, and so on. The gem uses locally running Redis as the default backend, and if that's not available, it uses the file system. The choice of the backend happens once and is saved in the git-ignored file in your local working repo.

Together, the three repos, after installation provide you with the consistent way to replicate your ~/.agents and ~/.claude folders on multiple computers, and to enable a consistent agentic software team workflow across any number of many projects.

Agentilda turns a short feature description into a reviewed pull request. You write the brief; a team of seven Claude Code agents researches it, specifies it, plans it, builds it and reviews it. You merge.

Install

gem install agentilda agent-lock -N
hash -r
tilda --help

agent-lock (the alock command) stops agents sharing a checkout from overwriting each other's files.

Both commands have tab completion. Add this to ~/.zshrc (or use bash in ~/.bashrc):

eval "$(tilda completion zsh)"
eval "$(alock completion zsh)"

help-screen

What it compiles

The dashboard is drawn by ratatui_ruby, a Rust extension, so installing needs a Rust toolchain (cargo) and clang with libclang on the machine. On a Mac, brew install rust and the Command Line Tools are enough; on Debian or Ubuntu, rustup plus clang libclang-dev. Without libclang, rb-sys's bindgen step stops at "Unable to find libclang", and gcc is not a substitute.

Building on a jemalloc-linked Ruby

If bundle install fails compiling ratatui_ruby with fatal error: 'jemalloc/jemalloc.h' file not found, your Ruby was built --with-jemalloc and rb-sys's bindgen step is not inheriting your compiler's include path:

BINDGEN_EXTRA_CLANG_ARGS="-I$(brew --prefix jemalloc)/include" bundle install

Put it in .envrc if you hit it more than once.

Quick start

Run these from the root of your project:

tilda create tax rule dsl   # creates .plans/000.00-⚪️ → tax-rule-dsl/spec.md and opens it
                            # finish the brief in spec.md, then save it
tilda run                   # dry run: shows which agent would take which plan
tilda run --commit          # runs the agents until no plan changes state

# alternatively
export AGENTILDA_AUTOCOMMIT=true 
tilda run 001.00 

tilda list-plans            # every plan, its state and its pull requests

The agents stop at an approved pull request. Merging it is up to you.

How it works

The folder name is the state

Every feature is one folder under .plans/, named NNN.MM-<emoji> → <slug>:

.plans/000.00-✅ → dev-foundation
       001.00-🟡 → tenancy-households
       001.01-🕰️ → schedule-k1-backfill      # shipped first, documented afterwards
       002.00-⚪️ → tax-rule-dsl
  • The number is set once and never changes. Branches, pull request titles and pull-requests.md all join on it. MM is 00 for a normal plan; 01 to 99 marks a retroactive plan slotted in after the fact.
  • The emoji is the state. A folder may only claim a state its files prove, such as spec.md, plan.md or pull-requests.md, and the tool refuses any other move.
  • The name contains spaces and emoji, so quote it in a shell.

tilda states draws every state and transition. docs/WORKFLOW.md, generated by tilda docs, is the full reference.

The lifecycle

stateDiagram-v2
    direction TB
    New: ⚪️ New
    Researched: 🔎 Researched
    Ready: 📋 Ready for Planning
    Planned: ⭐️ Planned
    Building: 🟡 Building
    BuildingUI: 🎨 Building UI
    Review: 🟢 Ready for Review
    InReview: 👀 In Review
    Changes: 🔴 Changes Requested
    Approved: ✅ Approved & Merged
    New --> Researched: leah
    Researched --> Ready: yoda
    Ready --> Planned: palpatine
    Planned --> Building: luke starts
    Building --> BuildingUI: luke done, rey building
    Building --> Review: last of the pair
    BuildingUI --> Review: last of the pair
    Review --> InReview: hansolo starts
    InReview --> Changes: rejected
    Changes --> Building: luke and rey fix it
    InReview --> Approved: a human merges
Loading

A plan can also leave the main path. It stops at ⭕️ Technical Block or 🅱️ Product Block when a human has to decide something. It ends at 💩 Scrapped by Review, ☢️ Deferred or ❌ Discarded when it won't be built.

The agents

Agent Takes plans at Moves them to Writes Job
leah-researcher ⚪️ 🔎 spec.md Researches the brief in parallel and appends a ## Research chapter
yoda-writer 🔎, 🕰️ 📋 spec.md Writes the full spec (goals, non-goals, scope), or blocked.md when a human must answer
palpatine-planner 📋 ⭐️ plan.md Splits the spec into independent work units, one plan each for back end and front end
luke-backend ⭐️, 🟡, 🔴 🟢 plan-backend.md, pull-requests.md Builds the data, domain, API and tests
rey-frontend ⭐️, 🟡, 🎨, 🔴 🟢 plan-frontend.md, pull-requests.md Builds the interface against Luke's API and proves the two halves work together
hansolo-reviewer 🟢, 👀 👀 approved, 🔴, 💩 pull-requests.md Reviews the diff against the plan; rejects at most twice; never merges
lando-broker ⭕️, 🅱️ ⭐️ plan.md Folds your answers from blocked.md back into the spec and plan

Luke and Rey work at the same time, in the same worktree, toward one pull request. They talk through mailbox.md in the plan folder, using tilda mail send and tilda mail read.

Run tilda agents list for the same table, or tilda describe <agent> to read one agent's prompt.

Unblocking a plan

The loop never assigns a blocked plan (⭕️ or 🅱️) to an agent. Write your answers into the plan's blocked.md, then:

tilda unblock 003            # what is 003 still waiting on?
tilda unblock 003 --commit   # fold the answered questions into spec.md and plan.md

Partial answers are fine. The plan stays blocked until the last question is answered.

Commands

tilda create tax rule dsl              # a new plan: 003.00-⚪️ → tax-rule-dsl
tilda create --from notes/dsl.md       # a new plan named by the file's frontmatter title
tilda create --after 002 k1 sync       # a retroactive plan: 002.01-🕰️ → k1-sync
tilda list-plans                       # every plan, its state and its pull requests
tilda index                            # write .plans/INDEX.md
tilda states                           # the state machine, as a diagram
tilda docs -o docs/WORKFLOW.md         # regenerate the conventions from the code
tilda run --commit --plan 003          # run the agents on specific plans
tilda unblock 003 --commit             # fold answered blocks back in
tilda worktree 003                     # create and seed a worktree for one plan
tilda resync dirs                      # rename folders whose emoji their contents contradict
tilda resync prs                       # add missing [NNN.MM] prefixes to pull request titles
tilda linear import TAX -p 'Tax DSL'   # mirror the plans into Linear

Important

Every command that writes is a dry run until you add --commit. Read what it would do first. Set AGENTILDA_AUTOCOMMIT=true to have every command behave as if --commit were passed.

Running the agents

tilda run                                  # dry run: who would take what
tilda run --commit                         # the whole tree, one worktree per plan, in parallel
tilda run --commit -j 4                    # at most four agents at once
tilda run --commit --plan 003,005.01       # only these plans
tilda run --commit --agent yoda-writer --prompt "Rework the risks section first"
tilda run --commit --skip hansolo-reviewer # everyone but the reviewer; its plans wait
tilda run --commit --model opus            # one model for every agent
tilda run --commit --timeout 600           # at most ten minutes per agent
tilda run --commit --scroll-height 5       # each agent's last five statuses under its row
tilda run --commit --agent luke-backend --prompt steer.md  # a short --prompt naming a file is read
tilda run --commit --rounds 1              # at most one round per agent per plan
tilda run --commit --max-tokens 200000     # token budget per agent invocation
tilda run --isolation shared               # one checkout, one agent at a time, no git needed

If you just created several plans, run them with a single --plan list. Separate bare run calls each loop over the whole tree, so their worktrees would overlap.

Below is the screenshot of multiple agents running, working on two separate plans. The fist plan has both Luke and Rey (backend and frontend developers) runnign, while the second plan is only at the research phase with Leah.

agents-running

The loop stops when a full round changes no plan's state. It also stops when every plan in scope is done, blocked, or out of rounds.

Safety rails

  • Isolation. By default every plan gets its own git worktree and branch, <user>/NNN.MM-slug, so agents on different plans share nothing. Parallel runs require it: --isolation shared runs one agent at a time.
  • Agents cannot publish. They may write source, tests and their plan's documents, but may not commit, push or touch a pull request. The tool enforces this twice: it withholds those commands from the agent, and afterwards checks that HEAD did not move.
  • The harness publishes. When a plan finishes, the harness commits the branch, pushes it and opens a pull request titled like [002.00](A) Tenancy Households. Pass --no-git-push to leave the work uncommitted in its worktree.
  • Nothing merges automatically. An approval leaves the plan at 👀 until you merge.
  • Disk over claims. A plan moves only when the files its new state requires exist. If an agent reports success but wrote nothing, the tool records it as interrupted.

Limits

Limit Declared by the agent Flag Default
Clock timeout: in its frontmatter --timeout 900 seconds
Rounds rounds: (at most 5) --rounds 1
Tokens none --max-tokens unmetered

Flags can only tighten what an agent declares, never loosen it. Each agent is told its clock and its token budget in its prompt.

As its time runs out, the agent is warned at 10, 5 and 1 minute left, then asked to stop. If it is still running 60 seconds after that, it is killed. A token budget works the same way: the agent is told the number, and the tool aborts it once it goes over.

Defaults for run can live in ~/.local/config/agentilda.json:

{ "run": { "timeout": 1800, "jobs": 4 } }

A typed flag beats the file, and the file beats the built-in default. The file accepts timeout, jobs, rounds, log and max_tokens. It never accepts commit.

The dashboard, and the keys during a run

On a terminal, every command that runs agents draws the same ratatui dashboard and listens for the same keys: run --commit, unblock --commit and create. A cyan status bar runs the full width at the top and another at the bottom, and each running agent is one table row — plan, agent, file, elapsed and remaining — with its latest statuses printed under it, newest first. run --scroll-height N (-s N) keeps the last N of them; the default is three. Off a terminal (a pipe, a dry run, cron) there is no screen, and progress goes to the log instead.

Key What it does
h show or hide this help
? about agentilda
s, ↓ select the next running agent; ↑ moves back
k mark the selected agent to be killed (STOP, 15 seconds, then kill -9)
x give the selected agent 10 more minutes, per press
ENTER apply the pending kills and extensions
ESC discard the pending changes, then clear the selection
w ask every running agent to wrap up as fast as possible
n ask agents to save their work and stop; the loop continues with the next agent
q save, stop everything and quit, after a 60-second grace period
ctrl-c agents leave a resume note and stop, then quit; press again to abort at once

k and x do nothing until you press ENTER, so a slip of the finger cannot kill an agent.

Ctrl-C once asks each running agent to write a RESUME: note into its plan's mailbox.md and sign Interrupted; every agent reads its mail before it starts, so a later run picks up where this one stopped instead of redoing the work. A second press aborts immediately. It works without a dashboard too, through a SIGINT handler.

How agents hand off

An agent never renames its plan folder. Instead, it signs the document it owns, and the harness moves the folder based on the signature:

> [!NOTE]
>
> [2026-09-04 11:29:20 AM PDT] [ agent: leah-researcher   status: Started, round 1 ]
> [2026-09-04 11:44:03 AM PDT] [ agent: leah-researcher   status: Completed, round 1 ]
> [2026-09-04 11:44:04 AM PDT] [ next: yoda-writer ]
Status What the harness does
Completed moves the folder on, if the files the next state requires are there
Blocked parks the folder at ⭕️, or at 🅱️ when the note says product
Almost completed, Interrupted gives the agent another round, up to its limit

Suppose an agent stops without signing, because it crashed, timed out or was killed. The harness signs Interrupted for it. If the work is on disk anyway, the harness also signs Completed for it. The run's own state, such as process ids and token counts, lives in .plans/agentilda-state.json, gitignored. A later run picks up where a dead one stopped.

GitHub and Linear

  • resync dirs renames folders whose emoji their contents contradict. For example, a ⚪️ with a plan.md becomes ⭐️. It is safe to run repeatedly.
  • resync prs adds the [NNN.MM] prefix to pull request titles. It works from the branch name first, then from the diff. It reports anything ambiguous and never changes it, even with --commit. It needs gh.
  • linear import TEAM creates one Linear issue per plan and one child issue per work unit. It only goes one way: nothing edited in Linear comes back. It uses LINEAR_API_KEY, or with --format json it hands the same data to the Linear MCP server.

From Claude Code

The /plan-create, /plan-run, /plan-status and other /plan-* slash commands wrap this same binary. They also add checks, such as confirming the scope and --commit before a run. They ship with agentilda-ai-setup. Use them inside a Claude Code session, and the binary everywhere else.

Development

just test         # rspec
just lint         # rubocop
just format       # rubocop -a, then mdformat
just ci           # lint, then tests with coverage
just update-workflow   # regenerate docs/WORKFLOW.md

The executables in exe/ find their own Gemfile, so they behave the same whether they are run from PATH or from another project's root.


© 2026 Konstantin Gredeskoul