Better Workflows is an open-source QA engineer and delivery gatekeeper for AI engineering — a demanding senior reviewer for every agent hand-off. A stage passes only when its evidence belongs to the current repository, revision, scope, and target, and can be checked again.
$$better-workflows:auto <describe the outcome you need>
›routeevidence-required · policy dev-publish-v1
›goalFreeze goal, scope, acceptance, and authority bound
›sourceBind the current repository and revision; take a source sentinel fresh
›worktreeCreate a task-owned branch and worktree; your checkout is untouched isolated
›executeRun bounded work inside the bound scope done
›verifyTyped evidence bound to the current source; review receipt present passed
›authorityThis target is authorized for a single side effect granted
›actPerform ONE external side effect sent
›reconcileReconcile provider state: returned outcome unknownunknownReconcile provider and repository state consistent
›completeNo retry, no completion claim not reachedRe-take the sentinel, re-verify acceptance, clean task-owned resources complete
STOPPED SAFELYProvider outcome unknown: no retry, no completion claim. The run stops and waits for your decision.
COMPLETETerminal provider and repository evidence is in place; the task-owned branch and worktree are cleaned up.
GATE STATE
STOPPED AT 08 / 09
Goal frozenGOALPASS
Source boundSOURCEPASS
Worktree isolatedWORKTREEPASS
Bounded executionEXECUTEPASS
Fresh, reviewed evidenceVERIFY · GATEPASS
Target authorizedAUTHORITY · GATEPASS
Single side effectACTPASS
Provider state reconciledRECONCILE · GATEBLOCKED
Complete and clean upCOMPLETENOT REACHED
PROVIDER OUTCOME
This is an illustrative walkthrough. Fields and states only show where gates pass and where the run stops; the alternative ending confirms the provider outcome and the run completes.
≥22.14Node.js versionminimum for the bundled helper
2Public languagesen · zh-Hant-TW
01PrinciplesPRINCIPLES
A prompt can describe intent. It never grants authority.
A stage passes only when its evidence belongs to the current repository, revision, scope, and target, and can be checked again. If evidence is missing, stale, conflicting, or the outcome is unknown, the workflow stops and asks you to decide instead of pretending the task is done.
TRUST 01
Root-owned mutation
Only Root may edit, integrate, deploy, accept risk, or declare completion. A prompt describes intent; it does not grant authority.
TRUST 02
Evidence before action
Every side effect requires fresh evidence, provenance, and an action bound to the intended target.
TRUST 03
Fail closed
Drift, stale evidence, or an unknown provider state always stops the workflow instead of pushing ahead.
WITHOUT vs WITHWithout governance vs. with Better Workflows
Aspect
Without governance
With Better Workflows
Authority
Intent and authority are conflated
Goal, scope, and authority are separate records
Freshness
A passing check may belong to an old revision
Evidence is bound to the current source and target
Retries
A retry may duplicate an external action
Attempts are bounded and unknown outcomes are reconciled
Done
“Done” can mean “the command returned”
Completion requires terminal provider and repository evidence
Isolation
Two tasks edit one checkout
Mutating Git tasks use separately owned branches and worktrees
AUTHORITY LAYERSAuthority layers
L1
Prompt
Records the outcome you want
Grants no authority
L2
Context
Binds current facts
L3
Harness
Limits who may act and where
L4
Loop
Bounds retries and reconciliation
L5
Graph
Projects admitted state; never a scheduler, policy input, or authority source
Projection only
02WorkflowWORKFLOW
Risk sets the verification, not ceremony.
A clear, reversible, low-risk change may use Auto's fast path with a small targeted check. Everything else is promoted to the evidence workflow, with verification strength matched to the risk.
AUTO FLOWAuto in five steps
Check the repository, goal, scope, and source branch
Choose Auto fast path or evidence-required
Read-only work stays in place; Git changes create or reuse an isolated task worktree
Validate the result; integrate changes only when authorized
Clean only task-owned branches and worktrees
ROUTINGFast path or evidence workflow
ENTRYPOINT$better-workflows:auto
BINDS ONE POLICYread-only-v1code-change-v1dev-publish-v1
AUTO DECIDESVerification strength follows the task and its risk
FAST PATH
Auto fast path
Clear, reversible, low-risk changes
A small, focused targeted check instead of the full evidence workflow.
Even on the fast path, it still
never bypasses a protected branch
never widens the scope
never installs tools
never skips the task-owned worktree
EVIDENCE WORKFLOW
Evidence workflow
Everything else
Verification strength matches the risk; evidence must belong to the current source and target.
These checks promote to evidence mode immediately
package-manager
network
child-process
native
checkout-external
LIFECYCLEFour questions replace “done”
01
Define
TaskContract
Are the goal, scope, acceptance, authority, and route frozen?
State the goal
Bind scope and current context
Git mutation?Yes ⇒ create or reuse a task-owned worktree
02
Verify
Evidence
Is the evidence bound to the current source?
Execute bounded work
Review and validate fresh evidencesource sentinel · typed evidence · graph and review receipts
03
Reconcile
Provider truth
Is the outcome of the external side effect confirmed?
Authorized for this target?No / unknown ⇒ stop safely
Perform ONE side effectsingle-use authority
Reconcile provider and repository stateUnknown ⇒ investigate, never blindly retry; stop safely
04
Complete
Terminal decision
After re-sampling, does acceptance still hold?
Re-take the sentinel; re-verify acceptance, ledger, review, and remote result
Complete and clean owned resources
Replay repeats the decision over the recorded evidence; it never repeats push, merge, deploy, or release.
GIT SAFETYGit boundaries
Changes happen in an owned worktree; integration uses a checked candidate and compare-and-swap.
G1
Read-only work stays in place.
G2
Mutating Git work always uses its own task branch and task-owned worktree — never your checkout.
G3
Integration uses a checked candidate and a compare-and-swap update.
G4
Only task-owned branches and worktrees are cleaned, and only with proof.
G5
Dirty state is never stashed or hidden.
G6
A clean, exclusive worktree created by the host is adopted instead of nesting another one.
03HostsHOSTS & PLATFORMS
Exactly where RC1 runs, stated plainly.
V5.0 RC1 covers Codex, Gemini CLI, and Qwen Code on macOS × Node.js 22/24 only. Claude Code, Linux, and Windows qualification is deferred to V5.1 and is not support today.
V5.0 RC1 public scope: hosts and operating systems
Host
macOS
Linux
Windows
CodexRecommended
RC1 public
V5.1 deferred
V5.1 deferred
Gemini CLI
RC1 public
V5.1 deferred
V5.1 deferred
Qwen Code
RC1 public
V5.1 deferred
V5.1 deferred
Claude Code
V5.1 deferred
V5.1 deferred
V5.1 deferred
NODE.JS 22/24 · bundled helper needs ≥ 22.14.0V5.1 deferred = qualification not finished; not support
Technical details: V4 historical support matrix (reference only)HOST-SUPPORT-V1
The V4 matrices below are historical. Current RC1 supports macOS with Codex, Gemini CLI and Qwen Code on Node 22/24; GA remains pending.
V4 · HOST-SUPPORT-V1
V4 AI / OS capability matrix
Evidence-first AI engineering QA and delivery gatekeeper.
Let AI agents choose verification strength by risk and finish work safely in an isolated environment.
Simple changes move fast; important work uses evidence gates; Git changes use a dedicated worktree by default.
native = host-native · core-bridge = shared control layer · unverified/unavailable are explicit limits.
AI host
task-contract
typed-evidence
replay
action-gate
task-worktree
native-picker
native-subagents
Codex
native
native
native
native
core-bridge
native
native
Claude Code
core-bridge
core-bridge
core-bridge
core-bridge
core-bridge
unavailable
unverified
Gemini CLI
core-bridge
core-bridge
core-bridge
core-bridge
core-bridge
unavailable
unverified
Qwen Code
core-bridge
core-bridge
core-bridge
core-bridge
core-bridge
unavailable
unverified
Kimi Code CLI
core-bridge
core-bridge
core-bridge
unverified
core-bridge
unavailable
unverified
Kiro
core-bridge
core-bridge
core-bridge
unverified
core-bridge
unavailable
unverified
Grok Build
core-bridge
core-bridge
core-bridge
unverified
core-bridge
unavailable
unverified
Cursor
core-bridge
core-bridge
core-bridge
unverified
core-bridge
unavailable
unverified
GitHub Copilot
core-bridge
core-bridge
core-bridge
unverified
core-bridge
unavailable
unverified
Official recommendation: macOS + Codex — the deepest native integration and complete reference experience.
Site source revision:16690d7c323e204b98efd3c3675e7933a66531e1
04InstallINSTALL
Pick your host. Paste the commands.
RC1 ships install paths for Codex, Gemini CLI, and Qwen Code on macOS. The bundled helper needs Node.js 22.14.0 or newer. Use a repository you trust; Better Workflows does not claim to sandbox malicious repository code.
1
Install
Choose your host and run the commands in a terminal. The official recommendation is macOS + Codex.
2
Reload
Codex: open a new task so the skill list refreshes. Gemini CLI and Qwen Code: restart the session after installing.
3
Make a first request
Type $better-workflows:auto, then describe the outcome you need.
RC1 is not GA. For every install step, see Quick start
›$better-workflows:auto Review this repository and fix verified defects.
›$better-workflows:auto Review this repository and summarize its main parts. Do not change files.
The first request fixes verified defects; the second is read-only.
05Proof boundaryPROOF BOUNDARY
What it proves, and what it doesn't — written down.
Up front: the errors it can block, what has not been proven yet, and what it is not.
Detects and blocks
Observable errors
The wrong repository or revision
Stale evidence
A false completion claim
An unauthorized side effect
An unknown provider outcome
Premature cleanup
Not proven yet
Long-term outcomes
Not statistically proven to lower scope drift, rework, or decision-error rates in multi-turn agent work
Cannot prove that your original goal was the right product decision
It is not
Outside the boundary
Not a sandbox: it does not claim to isolate malicious repository code — use a repository you trust
Not an unlimited agent runtime
It never harvests sensitive or private history
Better Workflows can block observable errors such as the wrong repository or revision, stale evidence, unauthorized side effects, and premature cleanup. It has not yet statistically proven lower long-term scope drift, rework, or decision-error rates.
B1Auto fast path still runs targeted checks and never bypasses protected branches
B2Protected or remote targets use governed PRs, fresh checks, and merge authority
B3Replay re-evaluates recorded evidence; it does not merge, push, or deploy again
06Status & licenseSTATUS & LICENSE
V5.0 RC1 is publicly available. GA remains pending.
RC1 is a controlled prerelease with Auto as the only public entrypoint. The GA conditions and deferred items are listed below — nothing is claimed early.
NOW Public since 2026-10-03
V5.0 RC1
5.0.0-rc.1 · V5.0.rc1
Codex, Gemini CLI, and Qwen Code on macOS × Node.js 22/24. Auto is the only public entrypoint.
PENDING Not released
GA 5.0.0
Requires all of
At least 30 natural canary days
20 consecutive eligible starts
Three distinct repositories
DEFERRED Deferred to V5.1
V5.1
Claude Code · Linux · Windows
Qualification for these hosts and operating systems is deferred to V5.1 and not yet released.
STATEMENTStatement of record
V5.0 RC1 · Public release and licensing
V5.0 RC1 (5.0.0-rc.1, tag V5.0.rc1) is a controlled prerelease with one public Auto entrypoint. The V4 support matrix remains historical; RC1 does not establish GA acceptance or V5 completion.
V5.0 RC1 covers Codex, Gemini CLI, and Qwen Code on macOS × Node 22/24. Claude Code, Linux, and Windows qualification is deferred to V5.1. GA requires at least 30 natural canary days, 20 consecutive eligible starts, and three distinct repositories.
The first-party Better Workflows core is AGPL-3.0-only. The physically separate minimal wire package is Apache-2.0; its LICENSE and NOTICE apply to that package.
The basic product is free. Professional Pack is planned as a proprietary product, and Cloud is a separate product planned for later; neither is currently available.
LICENSINGLicensing and availability
Licensing and availability
Item
License or form
Status
First-party core
AGPL-3.0-only
Public in RC1
Minimal wire package
Apache-2.0Physically separate; its LICENSE and NOTICE apply to that package
Public in RC1
Basic product
Free
Public in RC1
Professional Pack
Planned as proprietary
Not available
Cloud
Separate later product
Not available
07DocsDOCS
Read for your next step.
Five documentation pages in English and Traditional Chinese. To see the whole process first, start with Evidence Cinema.
An open-source QA engineer and delivery gatekeeper for AI engineering — a demanding senior reviewer for AI agents. A stage passes only when its evidence belongs to the current repository, revision, scope, and target and can be checked again; when evidence is missing, stale, conflicting, or unknown, the workflow stops and asks you to decide instead of pretending it is done.
Will it push, merge, or deploy on its own?
Not because a prompt said so. A prompt describes intent; it never grants authority. Only Root may edit, integrate, deploy, accept risk, or declare completion, and every side effect needs fresh evidence, provenance, and an action bound to the intended target — one side effect at a time.
Does every small change run the full process?
No. A clear, reversible, low-risk change can use Auto's fast path with a small focused check; everything else is promoted to the evidence workflow. Even the fast path never bypasses a protected branch, widens scope, installs tools, or skips the task-owned worktree.
Will it touch my current checkout?
Read-only work stays in place. Mutating Git work always uses its own task branch and task-owned worktree, never your checkout, and dirty state is never stashed or hidden.
What does RC1 support? What about Claude Code, Linux, and Windows?
The V5.0 RC1 public scope is macOS × Node.js 22/24 with Codex, Gemini CLI, or Qwen Code; the official recommendation is macOS + Codex. Claude Code, Linux, and Windows qualification is deferred to V5.1 and not yet released.
Does it guarantee bug-free code?
No. It can block observable errors such as the wrong repository or revision, stale evidence, unauthorized side effects, and premature cleanup, but it has not been statistically proven to lower scope drift, rework, or decision-error rates in long-running tasks, and it cannot prove your original goal was the right product decision.
Is it a sandbox?
No. It does not claim to isolate malicious repository code, so use a repository you trust. It is not an unlimited agent runtime, and it never harvests sensitive or private history.
What does it cost, and how is it licensed?
The first-party core is AGPL-3.0-only; the physically separate minimal wire package is Apache-2.0. The basic product is free. Professional Pack is planned as proprietary and Cloud is a separate later product; neither is available yet.
09SponsorSUPPORT
Help keep Better Workflows maintained.
Support with USDT (TRC20)
A one-time contribution supports open-source maintenance, English and Traditional Chinese documentation, and website hosting. It does not buy membership, roadmap priority, or support priority.
USDT · TRON (TRC20)
TGuMUi1d8MoBQcuFrGJZnu4JrbaeP3wy9a
V5.0 RC1 · publicly available
Make the next “done” something you can check.
Install the public RC1 and start with $better-workflows:auto. GA is still pending, and we will keep saying so up front.