jekyll-agent-audit
An offline audit of rendered Jekyll publications. It finds broken local references, conflicting document identity, structural markup problems, and observable metadata contradictions. It does not predict search rankings or AI citations, score pages, rewrite content, or check deployed HTTP behavior.
Install and run
Add the gem to your site's Gemfile, in the command-loading group:
group :jekyll_plugins do
gem "jekyll-agent-audit", "~> 0.1.0"
endFor an unpublished local checkout, add path: "/path/to/jekyll-agent-audit" to that declaration. Run bundle install, then:
bundle exec jekyll agent:audit
bundle exec jekyll agent:audit --format json --output tmp/agent-audit.json
bundle exec jekyll agent:audit --config _config.yml,_config.production.yml
bundle exec jekyll agent:audit --fail-on warning
bundle exec jekyll agent:audit --only links.target_missing,identity.canonical_multiple
bundle exec jekyll agent:audit --list-rules --format json
bundle exec jekyll agent:audit --explain identity.canonical_multiplePutting the gem only under _config.yml's plugins setting is too late for initial command discovery. See Jekyll's command integration documentation.
The command makes one full build in a temporary directory, with a separate temporary cache, and inspects the final output. Existing _site output is preserved. Build logs go to stderr; JSON stdout contains one report. --output writes atomically. Merely loading the gem does not audit an ordinary jekyll build.
The analyzer performs no network requests. Your site's ordinary build plugins retain their own behavior, including any network requests or custom filesystem writes they perform. Temporary build isolation cannot constrain arbitrary third-party plugin side effects.
Findings and exit codes
The default release enables these 17 rules:
| Severity | Rules |
|---|---|
| Error |
architecture.output_collision, links.target_missing, links.fragment_missing, identity.canonical_invalid, identity.canonical_multiple, identity.canonical_target_missing, structure.id_duplicate, provenance.jsonld_invalid
|
| Warning |
architecture.orphan, structure.title_missing, structure.heading_empty, structure.link_unnamed, provenance.date_invalid, provenance.date_order, provenance.metadata_conflict
|
| Info |
structure.h1_absent, structure.heading_jump
|
--explain RULE gives applicability, evidence, limitations, and remediation. Deferred catalog rules cannot be selected. Multiple H1s, an external canonical, no modification date, long prose, and no external citations are acceptable by themselves.
| Exit | Meaning |
|---|---|
| 0 | Completed; no active findings meet the failure policy |
| 1 | Completed; active errors, or warnings with --fail-on warning
|
| 2 | Invalid configuration/invocation, failed build/report write, resource limit, or incomplete supported-input analysis |
| 130 | Interrupted by SIGINT |
Default fail_on: error leaves warnings and info nonblocking. none disables finding-based failure; operational failures still return 2. A build with no HTML returns 2. Discovery always reports not assessed, owned by jekyll-agent-discovery.
Configuration
These are the defaults, under _config.yml. CLI audit options override them after ordinary Jekyll config layering. Jekyll environment, drafts, future posts, URL, and baseurl retain the site's configuration.
agent_audit:
fail_on: error
format: console
include: ["**/*"]
exclude: []
content:
selector: null
exclude_selectors: [nav, footer, aside, script, style, template]
metadata:
selectors:
author: null
publisher: null
published: null
modified: null
graph:
roots: ["/"]
routes:
directory_index: index.html
external_paths: []
aliases: {}
rules: {}
suppressions: []
limits:
max_html_bytes: 5242880
max_jsonld_bytes: 1048576
max_documents: 50000
max_edges: 1000000Unknown options, rules, invalid CSS selectors, alias cycles, and malformed limits are rejected. Disable a rule with rules: {structure.heading_jump: {enabled: false}}. --only intersects enabled rules and cannot re-enable one.
Include/exclude patterns match source-relative paths using Ruby File.fnmatch? pathname/glob semantics; exclusion wins. Excluded pages still supply links and targets to the graph. Generated HTML is included with a stable generated identity and a null source path when no source mapping is known.
Graph roots are site-relative: / becomes /project/ when baseurl: /project. Roots must exist. Route aliases and external_paths use full public paths, including baseurl. For example:
agent_audit:
routes:
external_paths: ["/project/api/**"]
aliases:
/project/old-guide/: /project/guide/These declarations describe hosting assumptions; they do not create or verify redirects. Assets are valid local targets. Directory-index aliases are proven by the emitted index file; /guide is not automatically equivalent to /guide/. Query strings stay in observations while target existence uses the path.
Content selection uses page agent_audit.content.selector, site selector, unique main, unique main-role landmark, then unique article. Body fallback supports whole-document checks and skips content-sensitive headings. Ambiguous selectors produce coverage diagnostics. Metadata date selectors read machine datetime attributes, never free-form prose.
Suppressions
Suppressions retain findings in JSON and require a reason. Site suppressions require source path scope and may specify a fingerprint. until is inclusive in the report's UTC reference date.
agent_audit:
suppressions:
- rule: architecture.orphan
paths: ["landing/private-preview.md"]
reason: "Linked only from a customer email."
until: "2026-12-31"Page front matter can use:
agent_audit:
suppress:
- rule: structure.h1_absent
reason: "The layout renders the title outside main."Expired suppressions do not hide findings. Unused suppressions generate notices. Collisions require a site suppression with primary-owner scope and a fingerprint. Disabling a rule appears as disabled coverage, never a clean result.
Reports and interpretation
Console and JSON share one report. The JSON schema is versioned 1.0; consumers should tolerate additive fields and use rule IDs instead of parsing message text. Findings preserve evidence level, observation, rendered location, related locations, remediation, fingerprint, and suppression. Source lines are null unless an actual source mapping exists; rendered HTML lines are not Markdown source lines.
The example JSON report was generated by a fixture with a broken local link, no H1, and an explicit modification date preceding publication.
Fingerprints exclude line numbers, temporary paths, message wording, and timestamps. Findings and coverage are deterministically ordered; run clocks and timings are volatile. Reports contain bounded observations, not article bodies. They can still reveal unpublished paths, so choose artifact access and retention accordingly.
For every rule, candidate = evaluated + not_applicable + skipped. Evaluated includes clean and finding-producing applications. Suppressing a finding does not change coverage. Counts describe local observations; there is no quality percentage.
Development and release status
bundle install
bundle exec rake test
bundle exec rake lint
gem build jekyll-agent-audit.gemspec
ruby script/benchmark.rbGenerated documentation
Generate Markdown API documentation and the LLM index from the Ruby source:
bundle exec rake docsThis replaces doc/ with fresh YARD Markdown, updates the documentation index in
doc/Jekyll/AgentAudit.md, and writes llms.txt with links relative to the gem
root. Both namespaces are indexed: the library under Jekyll::AgentAudit and the
Jekyll subcommand under Jekyll::Commands::AgentAudit. The generated files are
checked in and shipped with the gem; yard and yard-markdown stay development
dependencies. Authored guides belong in docs/, not the generated doc/.
For the individual steps, run bundle exec rake yard, then
ruby bin/generate_llm.rb.
To verify everything, regenerate documentation, and build the versioned gem and
SHA-256 checksum in pkg/, run:
bin/prepare_releaseIt stops at the first failure and never commits, tags, pushes, or publishes.
The implementation targets Jekyll 4.4.x. See verification and compatibility for combinations actually tested and benchmark measurements. Do not infer native GitHub Pages compatibility. The runtime HTML parser is explicitly declared as Nokogiri; no optional SEO, sitemap, or Markdown plugin is required.
See limitations and evidence before interpreting findings. The design specification in the parent workspace is the normative MVP catalog; deferred editorial and duplication analyses are not enabled.