hmd
Nothing ships unproven. Heimdall is an autonomous multi-agent orchestrator for Claude Code that turns one prompt into a finished project — and keeps moving the bar on what counts as done: every plan wires an external, falsifiable oracle so the implementation can never grade its own homework, and the merge stays blocked until the work is proven correct. Decomposes work, spawns parallel agents, enforces quality gates with CTO-level judgment. Live observability via a terminal-native HUD statusline and end-of-run summary card.
pinned to #dff5d28updated 2 weeks ago
Ask your AI client: “install plugins/hmd”.
Requires the metahub MCP server installed in your client. Set up MCP.
mh install plugins/hmdmetahub onboarded this repo on the author's behalf.
If you own github.com/randomittin/heimdall on GitHub, claim the listing to take over publishing. Your claim preserves the existing eval history and badges; only the curator label is replaced with verified-publisher on your next publish.
Stars
5
Last commit
2 weeks ago
Latest release
published
- #agent-orchestration
- #ai-agent
- #ai-coding
- #ai-orchestration
- #anthropic
- #automation
- #autonomous-agent
- #claude-agent
- #claude-ai
- #claude-code
- #claude-plugin
- #claude-skill
- #cli
- #coding-assistant
- #developer-tools
- #devtools
- #llm-agent
- #llm-framework
- #llm-tools
- #multi-agent
What's bundled
Items extracted from this plugin's manifest + directory tree.
Skills (4)
skills/designmatchUse when matching a React Native screen to a Claude Design HTML canonical at ≥95% visual parity on real Android/iOS hardware — closes the loop with a VQA stub mode, Playwright canonical renderer at…skills/heimdallAutonomous superskill manager — the main orchestrator agent for complex development tasks. Analyzes prompts, detects and assigns relevant skills, decomposes work into sub-projects, spawns parallel …skills/self-improveThe deliberate self-improvement loop for the maintainer — after every Nth maintenance cycle (or on /hmd:self-improve), step back from fixing issues and improve Heimdall's OWN capability. Ports karp…skills/stacks
Commands (8)
/autocommitCheck the current auto-commit state and toggle it:/autonomySet how much Heimdall does on its own before pausing for you./benchRuns `bin/heimdall-bench` — the documented entry door over the measurement/debloatTurn the H-2 bloat engine (`bin/bloat-gate`) on an ENTIRE existing repo. This is/demoRuns `bin/heimdall-demo` — a standalone, zero-model-cost shell script that gives a/designmatchAuto-runs `designmatch init <canonical-source> --app-dir <pwd>`:/dream`/dream` is Heimdall's OVERNIGHT self-improvement + maintainer pass. It does the deep/feedbackUse when a team dev wants to send feedback (friction, a request, a note on the
Subagents (8)
architectRead-only architecture and planning specialist. Analyzes codebase structure, identifies patterns and risks, designs solutions, and emits machine-readable plans (PLAN files + waves.json) with runnab…coderFeature implementation specialist. Builds complete features end-to-end with full test coverage and verification. Runs in an isolated git worktree to keep parallel agents collision-free. Use proacti…database-architectDatabase design agent. Schema modeling, migration strategy, query optimization, technology evaluation. Use when task involves data layer, models, or persistence.designUI/UX design agent. Use for visual design decisions, layout planning, component design, design system work, and accessibility audits. Integrates with design-for-ai skills when available.docs-writerDocumentation agent. Use for writing and updating documentation, README files, API docs, and keeping docs in sync with implementation changes.fixerBug fixer agent. Picks up open GitHub issues labeled 'bug' or 'seeker', creates a fix branch, implements the fix, runs tests, and raises a PR. Use for automated bug fixing from issue queue.heimdallAutonomous superskill manager. Use proactively for any multi-step development task. Decomposes work into sub-projects, spawns specialized agents in parallel, enforces quality gates, and maintains p…incident-responderProduction incident response agent. Root cause analysis, observability-driven debugging, rollback strategy, blameless postmortem. Use when production is broken or errors are spiking.
Hooks (1)
- git
Evaluation report
WarningsAutomated checks the publisher passed at publish time — structure, docs, safety, and whether the artifact behaves as claimed.dff5d28· 2 weeks ago
Maintenance
21Recent activity
last push today
Tests detected
1 test directory · 188 test files
CI configuration detectedwarn
no CI config found (looked for GitHub Actions, CircleCI, GitLab CI, etc.)
Add a simple workflow (lint + test on PR) — it tells consumers the artifact is built reproducibly.
Release history
1- releasecurrentdff5d28warn2 weeks ago
Contents
Heimdall 🛡️
A cloud bot that fixes your GitHub issues and opens a proven PR. You review, you merge.
Every PR ships the runnable evidence that the fix passes. The bot opens it on a heimdall/* branch as a scoped GitHub App — never as you, never on main, and it never self-merges. A human always gates the merge.
Get a bot PR on your repo
Once you've installed the Heimdall Maintainer GitHub App on your repo and run claude setup-token, it's two commands:
rr connect # registers your App install + captures your Claude cred
rr "fix the flaky test in payments and open a PR"
What happens: rr signs your task with your own Ed25519 key and enqueues it. A gated worker clones your repo with your team's Claude subscription and your GitHub App installation, runs the issue-resolution loop until the fix passes the gates, and opens a heimdall/* PR on your repo. You review it. You merge it.
Nothing to paste — no token, no URL. The public control plane is baked in and enrollment is automatic: your first signed call registers this device on first use. Just rr connect and go. (Running your own deployment, or need to re-gate enrollment behind a bootstrap token? That's an operator concern — see OPERATORS.md.)
Why it's safe to point at your repo
- Tenant isolation is a falsifiable oracle, not a promise. Every cross-tenant attack — IDOR by repo slug, cred read across teams, queue drain, installation-id swap, signed-request replay — has a named invariant and a red-line mutant test. Drop any gate and
test/heimdall-cp-authz-gate.test.shgoes red; the keystone suite passes only when every mutant is caught. Full invariant + attack matrix:docs/specs/2026-07-03-rr-isolation-invariants.md. - BYOC — no shared keys. You pay your own Claude tokens; your credential lands in your own per-team Secret Manager secret and is injected env-only into your job — never logged, never echoed, never readable by another tenant.
- Least-privilege bot. The App holds exactly Contents + Issues + Pull requests — no Administration, no Actions/Workflows, no merge capability. It can open a PR; it cannot touch branch protection or push to
main. - Honest bring-up. This loop was hardened over a live multi-tenant bring-up that shook out a run of production-only failures — Google's GFE rejecting GET-with-a-body, cold-start identity drift, jobs starving under scale-to-zero — each now documented as fixed in
deploy/cloud-run/README.mdand the runbook.
Under the hood — the verification orchestrator
The bot is powered by Heimdall's local engine: a Claude Code plugin that turns one prompt into finished, proven work. Every plan wires an external, falsifiable oracle so the implementation can never grade its own homework, and the merge stays blocked until the work is proven correct. Install it to run the same gates on your own machine.
Install
curl -fsSL https://raw.githubusercontent.com/randomittin/heimdall/v2.0.12/install.sh | bash
No sudo. No telemetry. Idempotent — re-run to upgrade. Reversible:
hmd uninstall # removes everything; nothing else was touched
Prefer to inspect first?
curl -fsSL https://raw.githubusercontent.com/randomittin/heimdall/v2.0.12/install.sh -o install.sh
less install.sh # function-wrapped, no eval, no base64 — what you read is what runs
bash install.sh
Prerequisites: Claude Code 1.0+ · Git · jq (brew install jq)
First run
hmd demo --run
Scaffolds a real full-stack task, builds it, ends with a summary card and a follow-up prompt. Safe to run sight-unseen — hmd demo (without --run) prints the plan and does nothing.
Why Heimdall
- Catches the silent failures — ordering races, whole-sequence invariants, missing subsystems that pass a naive green suite.
- Falsifiable gates — every gate is proven able to go red before it is trusted green. The corpus of real failure cases replays on every change; a regression that once shipped can never ship twice.
- Proof of correctness, not just generation — the delta Heimdall sells is the receipt that proves the proof can fail. Generalizes: 0.50 median reuse across 8 cold repos.
- Full audit trail —
hmd reportproduces a machine-readable telemetry report of every gate, mutation score, and corpus catch-rate from the last run.
What's inside
| Capability | Command | Status |
|---|---|---|
| Verification gates (secret-scan, bloat, falsify) | /hmd:verify in Claude · bin/falsify | Shipped |
| Demo task runner | hmd demo / hmd demo --run | Shipped |
| Issue-resolution loop | hmd (auto-retries failures against corpus) | Shipped |
| Telemetry report | hmd report | Shipped |
| Design match (visual diff vs spec) | hmd designmatch | Shipped |
| Redum / conformance checker | heimdall-redum · heimdall-check | Shipped |
| Reuse engine (cold-repo analysis) | bin/lib/reuse_analyzer.py | Shipped |
| Debloat scanner | heimdall-debloat --report-only | Shipped |
| Parallel workers | hmd --team N "task" (N tmux panes, independent — no shared state) | Shipped (no coordination layer) |
| Benchmark suite | heimdall-bench | Shipped |
Viral statusline — watchman, team wall, gate animation
Heimdall's status bar is a full-width, four-row watchman HUD. It renders entirely shell-side (zero model, zero context cost) and reads Claude Code's statusLine JSON on stdin. Three surfaces, by how far they spread:
Your sigil — the identity hook. Every Heimdall identity (HAID) gets a unique, deterministic pixel watchman: same identity, same sigil, forever. It anchors the left of the line, prints big on the install card, and shares as a postable block:
python3 sentinels/hmd-sigil.py --seed $HMD_HAID --size large # share/banner render
bash hooks/hmd-banner.sh --share # postable "my watchman" card
The seed is your HAID by default — automatic, stable, no PII in the art. Works solo on day one, before any teammate shows up.
The team watch wall — the headline, and the moat. When teammates also run hmd in the same repo, the bottom row becomes a live wall of their watchmen and what each agent is doing — gate state colored in, a teammate's cell flashing red the instant their gate denies. Nobody else can render this; it needs Heimdall's coordination substrate. The wall is empty until your team joins, so the feature recruits your team for you.
Presence is opt-in per repo — each running hmd heartbeats <repo>/.heimdall/team/<haid>.json (TTL ~30s; a stale file means the agent left). Names live in the repo's team dir and never leave it; default off for non-team repos. The watchman watches your gates, not your team. At squad scale the wall caps at the ~6 most-recently-active teammates plus a +N more tail so a wide terminal never wraps.
The deny flash — the clip. When a gate blocks, hmd-gate-anim.sh redraws the big watchman inline: a scanning pulse settling to a green sparkle on pass, or three red beats and ✗ BIFRÖST CLOSED on deny. TTY-only — in CI or a pipe it collapses to one clean final frame so logs stay readable.
bash sentinels/hmd-gate-anim.sh deny "oracle/falsify" $HMD_HAID
Wiring (settings.json):
{
"statusLine": {"type":"command","command":"bash ${CLAUDE_PLUGIN_ROOT}/hooks/statusline.sh"},
"subagentStatusLine": {"type":"command","command":"bash ${CLAUDE_PLUGIN_ROOT}/sentinels/hmd-subagent-statusline.sh"}
}
install.sh wires this for you — it registers both entries into your ~/.claude/settings.json (honoring $CLAUDE_CONFIG_DIR) using the absolute installed path, idempotently and without clobbering a statusLine you set yourself. The ${CLAUDE_PLUGIN_ROOT} form above is the plugin-hook spelling; a user-level statusLine resolves no such variable, so the installer fills in the resolved absolute path — which is why the HUD now reaches every dev, not just whoever hand-wired it in dev setup.
hooks/statusline.sh drives the full-width watchman and falls back to the legacy single line if python3 is missing — it never errors, never blocks. Already a ccstatusline (9.2k★) user? Keep your line and drop the watchman in as a Custom Command widget:
python3 sentinels/hmd-statusline.py --widget # just the watchman + verdict segment
The sigil ships solo-first (viral-cheap, no team required); the watch wall is the team-gated headline that lights up once presence is wired into your gate hooks.
Running on your own work
cd /path/to/your/project
heimdall --auto "build a real-time dashboard with auth and charts"
--auto runs a background safety classifier that blocks prompt injection and risky escalation. It is the default. --dangerously-skip-permissions exists but is not the default — only use it in a throwaway sandbox.
Failures visible on purpose
Live flagship status: evals/flagship/STATUS.md — the ❌ rows are kept in view. The corpus dip log and golden provenance are at evals/corpus/CORPUS-STATUS.md and evals/oracles/emulator-gb/fixtures/golden/VERIFICATION.md.
A verification system that can't show you its own failures can't be trusted with yours.
Contributing
- Stack packs (
skills/stacks/) — teach Heimdall a framework's conventions and build commands. - Oracle packs (
evals/oracles/) — add a falsifiable external gate for a new domain.
See CHANGELOG.md for release history.
License
Self-maintenance (auto-update + self-heal)
hmd keeps itself and its host current, in the background, on session start — both are throttled (~24h), detached (never block the session), idempotent, and opt-out:
- Plugin auto-update (
bin/heimdall-autoupdate): checks the installed version vs the latest GitHub release; if newer, re-runs the latest installer in the background (takes effect next launch; never hot-swaps the running session). Off:HEIMDALL_NO_AUTOUPDATE=1or~/.heimdall/no-autoupdate. - Claude Code self-heal (
bin/heimdall-cc-selfheal): on a NATIVE Claude Code install, auto-repairs the "✘ Auto-update failed" class — a stale npm-global@anthropic-ai/claude-codeconflicting with the native updater. It removes ONLY that conflicting package, ensuresautoUpdates:true, and re-runsclaude update. Never touches an npm/brew-managed install, never uninstalls anything else, never touches credentials. Off:HEIMDALL_NO_SELFHEAL=1or~/.heimdall/no-selfheal. Inspect:heimdall-cc-selfheal status.
Reviews
No reviews yet. Be the first.
Related
espalier-engineering
Train your AI coders the way you'd train a vine — discover your codebase's actual patterns, then encode them as rules, skills, agents, hooks, and a guided pipeline so generated code lands inside your conventions on the first try, not the fifth
before-you-build
Production-ready workflow orchestration with 91 marketplace plugins, 199 local specialized agents, and 162 local skills - optimized for granular installation and minimal token usage
cursor
Delegate implementation, web research, and codebase exploration to the Cursor CLI (headless agent -p). Part of cc-multi-cli-plugin. Requires the `multi` plugin.
mh install plugins/hmd