Skip to content

Troubleshooting

Most problems fall into one of three buckets: daemon not running, agent not found, or push not triggering the pipeline. This page walks each one.

First stop for anything: no-mistakes doctor.

Debug in this order

flowchart TD
  problem["Something is wrong"] --> doctor["Run no-mistakes doctor"]
  doctor --> daemon{"Daemon issue?"}
  daemon -- "yes" --> daemonpath["Check daemon status and daemon.log"]
  daemon -- "no" --> triggered{"Did the push trigger a run?"}
  triggered -- "no" --> gate["Check remote, hook, and socket"]
  triggered -- "yes" --> provider["Check agent or provider setup"]

That order matches the actual boundaries in the system:

  • local environment and binaries
  • daemon and gate wiring
  • provider-specific PR or CI integration

Daemon won’t start

Symptoms: no-mistakes daemon status shows stopped, or no-mistakes exits with “daemon not running.”

Start it manually

Terminal window
no-mistakes daemon start

This installs or refreshes the managed service (launchd, systemd user service, or Task Scheduler), then starts it. If service install or startup fails, it falls back to a detached daemon.

Check logs

Terminal window
tail -f ~/.no-mistakes/logs/daemon.log

Check for stale artifacts

A leftover socket from an unclean exit no longer blocks startup: the daemon probes the socket path before binding and removes it only when nothing is listening on it. A stale PID file can still confuse status reporting:

Terminal window
ls -la ~/.no-mistakes/daemon.pid ~/.no-mistakes/socket

If the PID file points at a process that’s no longer running, remove it and run no-mistakes daemon start again.

”a no-mistakes daemon is already running for this NM_HOME”

This error always means a genuinely live daemon: the lock it reports cannot go stale (see Daemon & Worktrees for the singleton-lock model). Manage that daemon with no-mistakes daemon status and no-mistakes daemon stop instead of deleting the lock file - deleting the file does not release the lock and only weakens the guard.

If the socket exists and the process is running but stuck or unresponsive, no-mistakes bounds the connection wait with daemon_connect_timeout and fails fast with an error naming the socket path instead of silently starting a second daemon. Restart the stuck daemon:

Terminal window
no-mistakes daemon stop
no-mistakes daemon start

If the socket file exists but nothing answers at all (a dead socket left behind by an unclean exit, e.g. a crash or SIGKILL), commands that ensure the daemon is running (no-mistakes, init, attach, rerun, axi run, axi respond) now fail fast with a connect to daemon socket error instead of silently starting a replacement daemon. The error message itself includes a (run 'no-mistakes daemon start' to recover) hint - run no-mistakes daemon start directly to recover, since it self-heals past a dead socket and starts a fresh daemon.

”configured worktree placement is unusable”

The daemon refuses to start while any worktree_roots entry names a directory it cannot create run worktrees in, and ~/.no-mistakes/logs/daemon.log names the offending entry. Because every command starts the daemon, that takes the whole CLI down until the entry is fixed: point it at a directory outside NM_HOME and outside every gated checkout, or remove it, then run no-mistakes daemon start.

Managed service logs

  • macOS (launchd): launchctl list | grep no-mistakes and check ~/Library/LaunchAgents/com.kunchenguid.no-mistakes.daemon.*.plist
  • Linux (systemd): systemctl --user status no-mistakes-daemon-* and journalctl --user -u no-mistakes-daemon-* -f
  • Windows (Task Scheduler): schtasks /query /tn "no-mistakes-daemon-*"

NM_HOME collisions

If you have multiple installs with different NM_HOME roots, each gets its own scoped service name (with a short suffix derived from the path). Make sure you’re looking at the right one - no-mistakes daemon status reports which.

no-mistakes update refuses or aborts

Symptom: update refuses because active pipeline runs are in progress, prompts because the daemon is running from a different executable path, or aborts because the daemon executable path cannot be determined.

update, daemon stop, and daemon restart all refuse by default while pipeline runs are active and list the affected runs; Daemon & Worktrees owns the guard’s exact rules, including why -y/--yes does not bypass it.

First inspect each listed run with no-mistakes axi status --run <id>. A parked CI gate can clear itself after its PR becomes terminal, including after a daemon restart. The ci_timeout reference owns the exact fail-closed reconciliation rules, and Daemon & Worktrees owns restart behavior.

After upgrading from an older release, starting the daemon automatically completes stale active rows that already have a persisted merged or closed PR state. Do not edit state.sqlite directly.

Only when you have confirmed it is acceptable for every remaining listed active run to fail, force the lifecycle operation:

Terminal window
no-mistakes daemon stop --force
no-mistakes update

Agent binary not detected

Symptom: doctor reports that gate validation is unavailable, or a run fails before its first pipeline step because no runnable agent was found.

This is a hard failure, not a degraded validation mode. no-mistakes will not silently skip review, test evidence, documentation, or agent-assisted lint and report the remaining work as a passed gate.

Check PATH

The daemon uses the same binary-discovery order described in Choosing an Agent. When it’s running through a managed service, it reloads PATH from your login shell on macOS and Linux and appends common install locations such as ~/.local/bin, ~/go/bin, ~/.cargo/bin, ~/bin, /opt/homebrew/bin, /usr/local/bin, /usr/bin, and /bin.

If a native agent is installed in a version-manager shim directory or another nonstandard location, set an explicit override in ~/.no-mistakes/config.yaml:

agent_path_override:
claude: /Users/you/.local/bin/claude

For agent: acp:<target> and ACP aliases such as agent: cursor, set acpx_path for the bridge. If the raw target command is also outside PATH, set its target key in acp_registry_overrides; agent_path_override applies only to native agents:

acpx_path: /Users/you/.local/bin/acpx
acp_registry_overrides:
cursor: /Users/you/.local/bin/cursor-agent acp

For Antigravity or Gemini-based driving agents, install a supported native agent CLI separately or configure a working ACP target such as agent: acp:gemini with acpx installed. The calling agent is the AXI driver, not an implicit pipeline-agent backend.

The daemon logs its effective PATH at startup in ~/.no-mistakes/logs/daemon.log with the message daemon environment ready. If the log contains login shell environment resolution failed or login shell environment resolution returned no entries, the daemon used a degraded fallback PATH that may omit version-manager directories such as nvm, fnm, or volta, so tools like pnpm may be missing.

Restart the daemon after installing a new agent

Terminal window
no-mistakes daemon stop
no-mistakes daemon start

Agents fail with “403 Request not allowed” behind a proxy

Symptom: runs fail and the step log shows agents (for example claude --print) unable to reach the network, often with 403 Request not allowed.

A managed daemon started by launchd or systemd inherits only a minimal environment, so it does not see the HTTP_PROXY / HTTPS_PROXY / NO_PROXY / ALL_PROXY variables from your shell. no-mistakes bakes any proxy variables that are set when you install or refresh the service into the generated service definition. If you set up the proxy after installing, re-run the installer or no-mistakes daemon restart (with the proxy variables exported) so they get baked in, then confirm them in ~/.config/systemd/user/no-mistakes-daemon-*.service on Linux or ~/Library/LaunchAgents/com.kunchenguid.no-mistakes.daemon.*.plist on macOS. Once baked in, the values survive later restarts and binary upgrades even from a shell that does not export them, so you only need the variables exported the first time. Windows Task Scheduler inherits your logon environment and needs no forwarding.

macOS App Management prompts during agent runs

Pipeline prompts steer agents to keep intentional writes inside the disposable worktree and avoid mutating system locations such as /Applications, Homebrew-managed packages, or global tool configuration. This reduces macOS App Management prompts from agent-invoked commands, but it is not an OS sandbox.

If you still see prompts, check the step log for commands that intentionally write outside the worktree and move that setup into your normal development environment or an explicit repo-local command. Requested test evidence may still be written under the managed evidence directory (<NM_HOME>/evidence/<run-id> by default). On GitHub.com/GHEC, supported image and video artifacts are uploaded when the PR is rendered; an orphan evidence branch is added when test.evidence.store_in_repo is enabled. The Global Config Reference lists the cases that leave a local citation instead. Normal tool temp or cache writes can still happen outside the worktree. Testing prompts ask agents to remove transient working-tree artifacts they created, such as downloaded models, caches, build outputs, large binaries, or generated data directories, before completion.

A pipeline step failed

Symptom: a run stops with a failed step.

Check the per-step log at ~/.no-mistakes/logs/<runID>/<step>.log. Fatal step errors are appended to that log, so failures such as rejected pushes include the returned error output there instead of only appearing in daemon.log.

Push fails with refusing to force-push

This means the live remote branch changed after the pipeline’s last observed head and contains commit(s) the validated worktree did not incorporate. no-mistakes refuses the push instead of overwriting that remote work.

Fetch and inspect the configured push target, then rebase or merge the remote work into your branch before pushing through no-mistakes again. If the overwrite is intentional, push manually to the actual remote after reviewing the commits that would be discarded.

Push fails with refusing to allow an OAuth App to create or update workflow ... without workflow scope

This means the branch touches a .github/workflows/*.yml or *.yaml file and the push credential (a GitHub OAuth token or PAT stored for the push target’s host) lacks the workflow scope. GitHub rejects the push before the pipeline can open or update the PR.

Resolve it by adding the workflow scope to your GitHub credential before pushing through no-mistakes again:

Terminal window
# If you authenticated gh via OAuth (web browser):
gh auth refresh -s workflow
# If you authenticated gh with a classic PAT, its scopes are immutable —
# create a new classic PAT that includes the workflow scope at
# https://github.com/settings/tokens, then re-authenticate:
gh auth login --with-token < new-pat.txt
# If you authenticated gh with a fine-grained PAT, its repository
# permissions are editable — set Workflows to Read and write at
# https://github.com/settings/personal-access-tokens (the token value
# stays the same, so no re-authentication is needed).
# Then configure git to use the refreshed credential:
gh auth setup-git

If your push target’s HTTPS remote embeds the PAT in its URL (for example https://<token>@github.com/...), gh auth setup-git updates only the credential helper — no-mistakes pushes using the token in the remote URL, so that URL must be refreshed too.

no-mistakes keeps its own copy of the push target’s URL on the gate’s bare repo, so updating the URL in your checkout alone is not enough: re-run no-mistakes init afterward so the gate picks up the refreshed URL.

Terminal window
git remote set-url origin https://<new-token>@github.com/<owner>/<repo>.git
no-mistakes init

If you push to a fork (see GitHub fork contributions), the fork URL is stored separately and a bare no-mistakes init preserves it. Pass the refreshed URL explicitly:

Terminal window
no-mistakes init --fork-url https://<new-token>@github.com/<fork-owner>/<repo>.git

Prefer authenticating through the credential helper (gh auth setup-git) over embedding a PAT in the URL — a clean URL with no embedded token needs no init after a credential refresh.

This only affects branches that modify workflow files. A branch that touches no .github/workflows/*.yml or *.yaml pushes normally with a standard repo-scoped token.

Rebase pauses because the branch carries unpushed default-branch commits

This means a local default branch ahead of origin/<default_branch> is a strict ancestor of your branch, so the branch may contain unrelated local-default work. no-mistakes pauses with an ask-user finding instead of silently bundling that ambiguous work into the PR. If the local default tip and your branch HEAD are equal, it treats the commits as the intended delivery work and continues.

Push the default branch to origin if those commits belong in the shared base, or rebuild the feature branch from origin/<default_branch> to remove the unrelated work before running the gate again. Approve the finding only when you have confirmed the local default-branch work belongs in the delivery branch.

git push no-mistakes doesn’t start a pipeline

Symptom: push succeeds but no-mistakes shows no active run.

Check the remote

Terminal window
git remote -v | grep no-mistakes

If it’s missing, run no-mistakes init again. Re-running init refreshes an existing gate and repairs the no-mistakes remote when it is missing. It also reattaches an existing gate after you rename or move the repo directory, as long as the old path no longer exists.

Check the receive hooks

The gate’s bare repo has a pre-receive hook that authorizes ref updates before mutation and a post-receive hook that notifies the daemon after an admitted push. Look at the gate path:

Terminal window
no-mistakes status
# gate path is shown in the output
ls -la <gate-path>/hooks/pre-receive <gate-path>/hooks/post-receive

Both hooks should be executable. If either is missing or non-executable, no-mistakes init will reinstall it for an existing no-mistakes-managed gate. For validated registered gates and strictly named legacy gates, no-mistakes daemon restart also installs missing no-mistakes-managed hooks and refreshes legacy managed hooks. An existing custom pre-receive hook is preserved behind the managed admission wrapper. Current managed hooks resolve the gate as an absolute bare-repo path before notifying the daemon, so a shell with a bad PWD value cannot accidentally report the gate as .. If notify-push.log mentions invalid gate path: ., refresh the managed hook with no-mistakes init or no-mistakes daemon restart, then push again.

Also check <gate-path>/notify-push.log. The hook now appends daemon notification failures there and prints the same error back to the pushing client.

Check the daemon socket

Both receive hooks talk to the daemon over ~/.no-mistakes/socket. If the daemon is not running, pre-receive admission fails closed and the push is rejected before any gate ref changes. Start the daemon and push again.

If the gate is older, re-running no-mistakes init or restarting the daemon also reapplies hook-path isolation when Git supports config --worktree. That protects the gate hook if a tool such as Husky wrote core.hookspath into shared git config from inside a linked worktree. Crash recovery owns the gate validation and migration rules used during restart.

PR step is skipped

Symptom: pipeline completes but the PR step shows skipped.

Check the Provider Integration requirements. Most common causes:

  • gh, glab, forgejo-axi, or tea not installed (or, for GitHub, not on PATH)
  • The provider CLI reports that it is not authenticated; on GitHub, a timed-out or interrupted gh auth status is reported separately from auth failure
  • Bitbucket env vars not set in the daemon’s environment
  • Upstream is not one of the hosts listed in Provider Integration
  • Self-hosted GitHub Enterprise on a hostname that is not github.com isn’t detected because gh isn’t configured for the host; run gh auth login --hostname your-ghe.example.com so detection finds it. Once detection succeeds, the availability check is host-scoped (gh auth status --hostname your-ghe.example.com), so a stale token on github.com or any other configured gh host can no longer falsely mark the GHE repo as unauthenticated.
  • Self-hosted GitLab on a hostname with no gitlab marker isn’t detected because glab isn’t configured for the host; run glab auth login --hostname your-gitlab.example.com so detection finds it. Once detection succeeds, the availability check is host-scoped (glab auth status --hostname your-gitlab.example.com), so a stale token on gitlab.com or any other configured glab host can no longer falsely mark the self-hosted repo as unauthenticated.
  • Self-hosted Gitea isn’t detected because tea has no login configured for the host; run tea logins add --url https://your-gitea.example.com --token <token> --name <name> so detection finds it. See Self-hosted Gitea.
  • A non-GitHub repo record has a fork URL set; fork MR/PR routing is currently GitHub-only
  • You pushed the PR base branch (PR step always skips there; this is the repository’s default branch, or the configured pr.base_branch when set)

CI step stuck or timed out

Symptom: CI step keeps monitoring an open PR longer than expected, or pauses after the idle timeout.

Monitoring while the PR remains open - even after checks are currently healthy - is intended behavior, because a later default-branch update can make the PR conflict or rerun CI. Once the CI monitor reports readiness and the PR is mergeable, the CI panel shows ✓ Checks passed and the terminal title switches to Checks passed, so you can tell when to go merge the PR; the signal clears automatically if checks start re-running or a new failure appears. A trusted no_ci: true declaration can establish readiness for a zero-check repository; an empty forge response without that declaration is not ready. The CI step reference owns the exact readiness and signal-clearing rules, including GitHub Actions runs that do not appear in the PR check rollup.

How long the monitor runs is controlled by ci_timeout in ~/.no-mistakes/config.yaml, an idle timeout that re-arms whenever the upstream default branch advances; the ci_timeout field reference owns the default, the unlimited keyword and its aliases, and the exact re-arm semantics. Older config files may still contain an explicit ci_timeout: "4h" value; update it if you want the newer default behavior.

If the PR is still open at the timeout, the step pauses for approval with findings for the open monitoring state or any known unresolved failures. You can approve, fix, or skip from the TUI or no-mistakes axi respond.

A park that happens before the timeout, with a finding that CI checks could not be read from the provider, means the check read itself is failing (after 6 consecutive failed polls, the step stops waiting instead of spinning to ci_timeout). The finding is provider-neutral and the step log shows the underlying provider error; for GitHub, a gh older than 2.50 rejects the gh pr checks --json call and needs upgrading. The same park on GitLab, Bitbucket Cloud, or Azure DevOps points at that provider’s CLI or credentials instead. Use no-mistakes axi abort only when you mean to cancel the whole active run.

Step looks quiet or wedged

Symptom: no-mistakes axi status shows an active step with last_activity prefixed by quiet, or a review/test/lint step appears to run for longer than expected.

quiet means the step has not recorded a step-log line or native-agent lifecycle event for longer than step_quiet_warning. It is only a liveness signal. It does not cancel the step, fail the run, or mean the pipeline is safe to bypass.

A quiet Review step still ends on its own: each fixer or reviewer invocation is independently bounded by review_agent_timeout, after which the run fails with a timeout diagnostic in the step log. This is an absolute wall-clock limit, not an activity-reset idle timer: an invocation that emitted output reports measured last-activity evidence, while a no-output invocation reports its measured no-output duration. step_quiet_warning remains status-only. A quiet Test step is bounded the same way by test_agent_timeout, covering the post-test evidence-gathering agent and a Test-repair turn. Every other agent-spawning step (Document, Lint, Rebase conflict repair, PR drafting, CI auto-fix) is bounded by agent_timeout, so a stall reaches the step’s normal agent-error handling instead of remaining active until you abort. Most mutation steps fail, PR drafting continues with deterministic fallback content, and CI auto-fix parks for a user decision as described in the CI step reference.

Start by reading the active run and the step log:

Terminal window
no-mistakes axi status
no-mistakes axi logs --step <step> --full

See the axi status reference for active-step timing, activity, PID, and round fields. The step log records native subprocess start, exit, and retry lines plus markers for automatic and user-triggered fix rounds. If the step is parked at a gate, use no-mistakes axi respond instead of waiting. If the run is genuinely stuck and you want to discard it, use no-mistakes axi abort. Start a new run only after abort confirms the terminal state; see the abort command contract.

Worktree won’t clean up

Symptom: ~/.no-mistakes/worktrees/<repoID>/<runID>/ - or <root>/<runID> when the repository has a configured worktree root - sticks around after a run ends.

The daemon’s retention rules and crash-recovery checks can deliberately keep a worktree after a run ends. Inspect retained work before considering removal; for a protected-path refusal, follow the resolution guidance. Only remove a leftover after deciding its contents can be discarded:

Terminal window
# From inside the repo the worktree belongs to:
git worktree list
git worktree remove --force <path>

Otherwise, eligible orphan worktrees are cleaned on the next startup, subject to those same retention rules:

Terminal window
no-mistakes daemon stop
no-mistakes daemon start

Reset everything

When state is genuinely wedged:

Terminal window
no-mistakes daemon stop --force
rm -rf ~/.no-mistakes/worktrees ~/.no-mistakes/servers ~/.no-mistakes/socket ~/.no-mistakes/daemon.pid ~/.no-mistakes/daemon.lock
no-mistakes daemon start

If worktree_roots places a repository’s runs outside NM_HOME, delete the leftover <root>/<run id> directories there as well - only those; the root is your own directory and holds nothing else of no-mistakes’.

This keeps your gate repos, database, and config but clears transient state. For a full wipe, see the Uninstall section. Wedged state often means a run is stuck pending or running, so daemon stop refuses without --force; only force through once you’ve confirmed it’s fine for the listed runs to fail.

Still stuck