Blog

Agent of the Day – August 28, 2026

Most teams write lint rules once and trust them forever. Nobody goes back to ask whether the rule’s own error message is still telling the truth, or whether a fix applied to one rule ever made it to its dozen siblings. Today’s spotlight, ESLint Refiner, exists precisely for that blind spot: a daily workflow that treats gh-aw’s custom ESLint rule set — eslint-factory — as a codebase worth auditing in its own right, not just a tool you point at other code.

eslint-factory is the internal library of custom lint rules that keep actions/setup/js scripts safe — things like “wrap this fs.mkdtempSync call in a try/catch” or “don’t compare objects with JSON.stringify equality.” Rules like these accumulate fast; a sibling workflow, eslint-miner, adds roughly one new rule a day. But nobody was reviewing whether the rules themselves stayed correct as the pile grew. ESLint Refiner picks two of the least-scrutinized rules each run, checks their logic against real call sites in the codebase, and only files an issue when it finds something grounded in an actual bug — not a hypothetical.

In its August 27 run, the agent reviewed require-mkdtempsync-try-catch and require-decodeuricomponent-try-catch, and turned up two real problems:

  1. A recurring overclaim. Eleven rules — including both reviewed that day — say a call “will crash the action if unhandled.” That’s not quite true: every entrypoint in actions/setup/js already has a top-level try/catch that routes any uncaught throw into a controlled core.setFailed, so nothing actually crashes silently. The real cost of skipping the fix is losing a specific { cause } and message, not crashing. This exact wording had already been fixed once, for require-fetch-response-body-try-catch, but the fix never propagated — and had since recurred in two brand-new rules. The agent filed issue #56288 asking for a reword-all pass plus a guard so it can’t quietly recur a third time.
  2. A misclassified literal. require-decodeuricomponent-try-catch only recognized string literals as provably safe arguments, so calls like decodeURIComponent(42) or decodeURIComponent(null) got flagged even though a number, boolean, or null can never produce a decoding error. Zero live call sites hit this today, but it’s a cheap, well-scoped fix worth closing before one does — filed as issue #56289.

Rather than silently move on, the agent also published a same-day discussion post summarizing exactly what it checked, what came back clean, and what’s queued for tomorrow’s review — including a note that its own repo-memory had gone stale for weeks even while it kept filing issues, which it then rebuilt from a ground-truth GitHub search.

The following day’s run, on August 28, kept the streak going with another clean pass — no errors, three safe outputs produced, business as usual for a workflow that’s quietly been doing this every day since it launched.

What makes ESLint Refiner worth spotlighting isn’t flashy output — it’s discipline. It doesn’t flag speculative issues; it grounds every finding in live call sites before filing, and it explicitly tracks precedent (citing the earlier fix for the same defect class) so fixes actually propagate instead of getting re-invented rule by rule.


Curious how workflows like ESLint Refiner are built? Explore the project at github.com/github/gh-aw.

Agent of the Day – August 26, 2026

Agent of the Day – August 26, 2026: The CLI Archivist

Section titled “Agent of the Day – August 26, 2026: The CLI Archivist”

Every CLI accumulates small inconsistencies over time — a flag renamed here, a doc page that forgot to follow, an extra blank line nobody meant to leave in. Today’s spotlight workflow exists purely to hunt down that kind of drift before a human ever notices it: the CLI Consistency Checker.

This is a patient, unglamorous job done exceptionally well. Each day the workflow collects the full --help output for every one of the gh aw CLI’s 40-plus top-level commands and subcommands, then lines it up against docs/src/content/docs/setup/cli.md looking for typos, flag-naming inconsistencies, missing --no-* negation counterparts, undocumented commands, and stale examples. It’s the kind of exhaustive line-by-line comparison a person would dread doing manually — which is exactly why it’s automated.

Two consecutive runs this week show the pattern at its best. On run 32924017960 (August 25), the checker filed issue #55788, flagging that the graders command — a fully functional first-class CLI command with its own operational-value subcommand — was completely absent from the documentation. That report was closed out same-day by PR #55794, which added the missing docs.

The next day’s run, 32974417489, turned up something subtler: issue #56047 reported that seven commands — including gh aw add, gh aw logs, gh aw trial, and three mcp subcommands — were rendering two blank lines before their Flags: section instead of the single blank line used everywhere else. The root cause traced back to trailing newlines left inside Go raw string literals for each command’s Example: field. Cosmetic, yes, but the kind of thing that makes a CLI feel polished versus slightly off. PR #56052 trimmed the stray newlines and merged the same day, restoring consistent formatting across all seven commands.

What stands out across both runs is the discipline of the reports themselves: a clear severity breakdown, an affected-commands table, a root-cause section pointing at the exact source file, and — critically — zero false positives. Neither run flagged noise; every finding led directly to a merged fix. That’s the bar an agent needs to clear to earn trust running unattended on a schedule.

It’s a quiet workflow — no flashy dashboards, no dramatic incident response — just a steady daily diff between what the CLI says it does and what the docs say it does. But drift like this compounds silently in any fast-moving codebase, and catching it same-day, every day, is exactly the kind of tedious vigilance agentic workflows are built for.

Curious how a workflow like this is defined? Check out github/gh-aw and see how a few lines of markdown frontmatter turn into a disciplined daily CLI audit.

Agent of the Day – August 25, 2026

Section titled “Agent of the Day – August 25, 2026: The Cookie Monster of Issues”

Some workflows in gh-aw wait patiently for a human to summon them. Today’s spotlight isn’t one of those. It wakes up every 30 minutes, peers into the open issue tracker, picks out the tastiest morsel it can find, and hands it straight to the Copilot coding agent. Its name says it all: Issue Monster.

Described in its own frontmatter as “the Cookie Monster of issues,” this workflow runs on a schedule: every 30m trigger with a set of guardrails that keep it from overreacting. It skips a run entirely if there are already five or more open draft PRs from app/copilot-swe-agent, skips if there are no open issues to consider, and skips if key CI checks (build, test, lint-go, lint-js) are failing. Before picking a new target, it even checks for recent rate-limiting signals on Copilot-authored PRs from the last hour, so it doesn’t pile more work onto an agent that’s already struggling.

Five real runs from the last day tell a consistent story of steady, careful triage:

  • Run 32853285457 (8.9 minutes, 3 turns) evaluated the tracker and moved on without a strong enough candidate that cycle.
  • Run 32856106356 found three good bites in one pass, assigning issue #55788, issue #55770, and issue #55716 to the Copilot coding agent, each with a HIGH-confidence rationale that they were “clearly scoped, independent candidates for automated resolution.” Every assignment came with a cheerful comment: ” Issue Monster selected this for Copilot… Om nom nom! ”
  • Run 32859101860 picked up the pace again shortly after, assigning issue #55771 and issue #55768.
  • Run 32821436131 and Run 32816269740 ran earlier in the day, each completing in 6–8 minutes with clean, error-free conclusions.

Across all five runs, the workflow burned through 350k tokens and racked up 125 GitHub API calls — mostly reads, scanning issue bodies, checking recent PR activity, and verifying rate-limit safety before ever touching the assign_to_agent safe output. Of the five runs, two executed write-capable safe outputs (the assignments and comments above) while three stayed strictly read-only, quietly confirming there was nothing worth biting into that cycle. Zero errors, zero warnings, across the board.

That restraint is the real design story here. Issue Monster only has issues: read and pull-requests: read permissions directly — it never edits code itself. Its entire job is curation: reading the room, checking capacity, and making a narrow, well-reasoned call about which issue is ready for an autonomous fix versus which one still needs a human’s judgment. The assign_to_agent safe output does the heavy lifting of actually routing the issue to Copilot, and the accompanying comment keeps the paper trail visible to anyone watching the issue.

gh-aw workflow activity chart

It’s a small, unglamorous job — but multiplied across dozens of runs a day, it’s the difference between an issue tracker that silently accumulates backlog and one where fixable problems get routed to an agent within half an hour of becoming “clearly scoped.” Om nom nom, indeed.

Curious how Issue Monster — or any other gh-aw workflow — is put together? Explore the project at github.com/github/gh-aw.

Agent of the Day – August 24, 2026

Agent of the Day – August 24, 2026: The Workflow Doctor

Section titled “Agent of the Day – August 24, 2026: The Workflow Doctor”

Most agentic workflows in gh-aw run on a timer, quietly doing their thing every day whether anyone’s watching or not. Today’s spotlight is different: it only shows up when you call it. Type /q in a comment on an issue, pull request, or discussion, and this workflow wakes up, reads the room, and goes to work fixing whatever you pointed it at.

We’re calling this persona The Workflow Doctor, and it belongs to Q, a slash-command-triggered gh-aw workflow described in its own frontmatter as an “intelligent assistant that answers questions, analyzes repositories, and can create PRs for workflow optimizations.” Q runs on the Copilot engine with SDK mode enabled, has read access to issues, pull requests, and discussions, and — critically — is under a hard rule never to touch its own definition file (q.md). It exists purely to diagnose and improve other workflows in the repo.

Three real runs from the last few days show exactly how varied its case load gets:

  • Run 32726221560 fired from a comment on discussion #55296, where a maintainer asked Q to “add a job that tests the github MCP in remote mode without using any agentic workflow feature” as a canary test to rule out a runtime/compiler bug, plus a summary of the MCP handshake message. Q completed in 11.5 minutes across a single turn, burning 25.8k tokens, and wrapped up with a successful conclusion and a proposed pull request queued up.
  • Run 32727642327 answered a comment on issue #55389 asking Q to “use mai flash model to reduce cost” — a straightforward cost-tuning request that Q turned into a workflow-level model swap.
  • Run 32726004596 came from discussion #55334, where the ask was to “update to use repo-memory to store the mined loops” — plumbing persistent state into a workflow that was previously stateless between runs.

Audit data on the discussion-triggered run classifies its behavior fingerprint as directed execution with narrow tool breadth and a selective_write actuation style — in plain terms, Q doesn’t wander. It reads exactly what it needs (the triggering comment, the parent issue or discussion, recent logs and audits for the target workflow), forms a specific diagnosis, and proposes a scoped pull request through its create-pull-request safe output, complete with a [q] title prefix, automation and workflow-optimization labels, and Copilot as the default reviewer.

That safe-outputs configuration is worth calling out on its own: PRs expire after 2 days if unmerged, patches are capped at 500 files, and protected-file edits automatically fall back to filing an issue instead of silently failing. It’s a small but deliberate guardrail set for a workflow that has write access to propose changes across the entire repo’s workflow surface — tight enough to keep blast radius small, generous enough to let Q actually fix things.

The interesting part isn’t any single fix — it’s the range. In three runs pulled from the same short window, Q handled a low-level infrastructure canary test, a cost-optimization tweak, and a state-persistence upgrade, each triggered by a different person from a different corner of the repository. That’s the value proposition of an on-demand workflow doctor: no scheduling, no queue, just /q and a clear ask.

Want to see how Q — or any other gh-aw workflow — is built? Explore the project at github.com/github/gh-aw.

Weekly Update – August 24, 2026

Another busy week for github/gh-aw! We shipped three pre-releases (v0.87.1, v0.87.2, and v0.87.4) and merged dozens of pull requests focused on safe-output reliability, compiler robustness, and internal reporting accuracy. Here’s what stood out.

The v0.87.4 line of releases focused on compiler robustness, safe-output validation, and internal tooling and observability improvements across the agentic workflow pipeline.

  • Run steering (#55171): safe-outputs.create-pull-request.pre-create.steer: true introduced run-scoped feedback issues and injected prompting for agents to read comments containing the steer keyword. The configuration was subsequently moved to safe-outputs.steer, where it requires explicit issues: read without silently expanding workflow permissions.
  • gh aw models (#55148): a new CLI command surfaces catalog pricing, alias resolution, and observed automation models in one place.
  • Copilot SDK startup diagnostics (#55149): pre-ready crashes now surface the Copilot SDK’s startup stderr, making a previously opaque failure mode much easier to debug.
  • Automatic PR review dismissal ingestion (#55180): workflows can now ingest automatic pull request review dismissals as part of their safe-output processing.

The tireless kitchen staff of the PR pipeline — it watches over open pull requests and keeps their descriptions, context, and footers fresh and useful.

This week pr-sous-chef ran three times in a single day (all successful, all on the pi engine), burning through roughly 76K tokens and racking up 8 safe-output items across its runs — including landing the new pre-create PR steering feature itself via #55171. Its most recent run alone produced 5 safe items in under 8 minutes, a tidy little burst of productivity right before this post went out.

Somewhat fittingly, the workflow that teaches other PRs how to listen to reviewer feedback (steer) shipped that very feature about itself — a small bit of “eating your own dog food” that we appreciated.

Usage tip: Reach for a pr-sous-chef-style workflow whenever your team’s biggest bottleneck is PR descriptions and footers going stale between review rounds — it keeps that metadata current without anyone lifting a finger.

→ View the workflow on GitHub

Grab the latest v0.87.4 release to try the initial pre-create PR steering configuration. Current workflows use safe-outputs.steer as described in #55792. As always, questions, bug reports, and contributions are welcome over at github/gh-aw.