Blog

Weekly Update – September 7, 2026

Another busy week for github/gh-aw! The team shipped a new release focused on hardening the agentic firewall and CI reliability, while dozens of pull requests tightened up sandboxing, model configuration, and safe-output handling across the fleet of agentic workflows.

v0.88.4 landed this week, focused on hardening the agentic firewall/network layer, improving CI reliability, and expanding project tooling support.

  • Trusted enclave sensitivity support (#58328): adds finer-grained sensitivity controls for trusted enclave workflows.
  • DIFC policy generation for GitHub App workflows (#58302): automatically generates data-flow integrity/confidentiality policies for workflows authenticated via a GitHub App.
  • aw.json project support in the add command (#58267): makes it easier to add workflows into existing aw.json-based projects.
  • Daily Linear and Jira smoke issues workflow (#58320): adds scheduled smoke-testing coverage for Linear and Jira integrations.
  • Fixed root-relative workflow paths in imported local manifests (#58317).
  • Fixed Pi Anthropic routing through the firewall (#58313).
  • Disabled OTLP export when authorization secrets are empty, avoiding noisy failed exports (#58312).
  • Preserved setup-ruby PATH precedence inside the Agentic Workflow Firewall (#58311).

Beyond the release, the past week’s merge queue was dominated by fleet-wide reliability work:

The tidiest member of the team — it scans the Go codebase for unreachable functions using static analysis and opens a PR to remove a small batch every day.

This week dead-code-remover ran three times and kept up its steady rhythm: two clean runs each produced a PR titled “[dead-code] chore: remove dead functions — 5 functions removed” (#58996, #58822), quietly trimming five functions each time, while one run hit a snag and came back empty-handed rather than force through a bad batch.

It never removes more than five functions per run — a self-imposed diet that keeps every PR small enough for a human to review in a coffee break, and disciplined enough that nobody’s ever caught it trying to sneak in a sixth.

Usage tip: Cap batch size like this for any “cleanup” agent — small, reviewable PRs land far more often than one giant sweep.

→ View the workflow on GitHub

Update to v0.88.4 and check out the new trusted enclave and DIFC policy features. As always, feedback and contributions are welcome in github/gh-aw.

MicroVM Support Is Consolidating on Cloud Hypervisor

GitHub Agentic Workflows has consolidated its specialized sandbox runtime support on cloud-hypervisor. The gvisor and docker-sbx runtime options have been removed.

docker-sbx introduced a KVM-backed microVM boundary, while gvisor provided a user-space kernel between the agent container and host kernel. Maintaining both paths alongside Cloud Hypervisor created separate installation, compatibility, and troubleshooting surfaces. Consolidating on one microVM implementation makes the stronger isolation path more consistent and easier to evolve.

The default docker runtime remains available and continues to run AWF with network isolation and proxy enforcement. For workflows that require a hardware-virtualized boundary, cloud-hypervisor is now the supported direction:

---
on: issues
sandbox:
agent:
runtime: cloud-hypervisor
---
Investigate this issue.

Cloud Hypervisor support is currently in preview and requires a GitHub-hosted Ubuntu x86_64 runner with /dev/kvm. The compiler adds the required host checks and provisions digest-pinned runtime assets.

Review workflows that explicitly set runtime: gvisor or runtime: docker-sbx. Select cloud-hypervisor when the workflow runs on an eligible GitHub-hosted runner and needs a microVM boundary. Otherwise, remove the runtime setting to use the default Docker profile.

Compile each updated workflow and review the generated lock file:

Terminal window
gh aw compile

The deprecated values remain documented during the transition, but new workflows should use either the default Docker runtime or cloud-hypervisor. See Agent Runtime Selection for requirements and tradeoffs.

Agent of the Day – September 3, 2026

Some workflows in gh-aw wait for a slash command or a pull request event before they lift a finger. Today’s spotlight has no patience for that. Every 40 minutes or so, it wakes up on its own schedule, scans the open issue tracker, picks out the most promising candidates, and hands them straight to the Copilot coding agent. Its name is exactly as subtle as its appetite: Issue Monster.

Issue Monster is a scheduled gh-aw workflow (.github/workflows/issue-monster.md) whose entire job is triage-by-appetite. It has a pre-fetched view of the open issue queue, a short list of skills — issue-monster-report-formatting and issue-monster-token-budget — to keep its output tight and cheap, and exactly three safe-output moves per run: assign_to_agent, add_comment, and noop if nothing qualifies.

Watching a stretch of its actual runs from earlier today tells the story better than any spec could. Between run #33740299445 and run #33766461931 — ten consecutive runs spanning about four hours — Issue Monster completed successfully every single time, burning roughly 1 million tokens and 103 action-minutes total, with zero errors, zero warnings, and zero missing tools across the board.

What makes the run interesting isn’t the token count, it’s the reasoning trail. In its most recent run, the agent’s internal notes show it weighing overlapping candidates before committing:

“I see some top candidates like #57408 (domains audit), #57728 (stale code-scanning alert), #57709 (CLI consistency)… I must ensure these issues are distinct and not overlapping… I’ll aim for the highest-scored independent issues.”

It settled on three genuinely separate issues — a security-hardening fix, a documentation gap, and a refactor — and for each one called assign_to_agent followed immediately by a comment announcing the decision:

Issue Monster selected this for Copilot — I’ve identified this issue as a good candidate for automated resolution and requested assignment to the Copilot coding agent. Om nom nom!

The cookie-monster signature isn’t just flavor text; it’s a consistent, greppable marker across issue #57408 (default domain allowlist hardening), issue #57709 (CLI docs consistency), and issue #58148 (workflow-skill extraction refactor) — three issues that, as of this run, are now sitting in the Copilot coding agent’s queue waiting for a PR.

Earlier runs in the same window picked up other issues from the same rotating shortlist, including #57142 (a Go package refactor) and #58240 (a large parser file split), showing the agent isn’t just repeating the same three picks — it re-evaluates the queue and reprioritizes as issues get claimed or closed.

The gh-aw audit tooling flagged one low-severity note worth mentioning: Issue Monster’s task profile (narrow tool breadth, read-mostly posture, moderate resource use) is a candidate for a cheaper model like gpt-4.1-mini instead of a frontier engine — a reminder that even a workflow running dozens of times a day has room to trim its own cost curve. That’s the kind of self-aware feedback loop gh-aw’s observability tooling is built to surface automatically, run after run.

No fanfare, no dashboard to babysit — just a small, disciplined loop that turns “here’s an open issue” into “here’s an assigned Copilot task” every 40 minutes, day and night.


Want to build something with the same rhythm? Explore the workflows, skills, and safe-output patterns behind Issue Monster at github.com/github/gh-aw.

Agent of the Day – September 2, 2026

Agent of the Day – September 2, 2026: The Complexity Cop

Section titled “Agent of the Day – September 2, 2026: The Complexity Cop”

Most PR bots check style, tests, or security. Today’s spotlight, Ponytail Reviewer, checks something harder to quantify: whether a change is more complicated than it needs to be. It runs on every pull request marked ready for review in gh-aw (and on demand via a /ponytail slash command), applies the community-maintained ponytail-review skill, and only speaks up when it finds real over-engineering — no noise, no rubber-stamping.

Pulling the last three runs from agentic-workflows logs and audits paints a workflow that’s doing its job quietly and reliably:

  • Run #33596102316 — a 7-minute pass over PR #57860 (branch copilot/sergo-fix-linters-silent-delete), completed successfully with the Codex engine, burning 164K tokens across 8 model requests.
  • Run #33637292011 — another clean 7-minute review, this time against a copilot/task-9919-* branch, also completing without incident.
  • Run #33635177913 — a quick 7-second run against a PR titled “Fix daily-token-consumption-report: replace unsupported claude-sonnet-4.5 model”, which failed fast rather than burning minutes on a doomed invocation — exactly the kind of fail-cheap behavior you want from an automated reviewer.

Across all three runs: zero errors, zero missing tools, and two safe-output items generated in total — meaning Ponytail Reviewer isn’t just running, it’s making judgment calls about when a comment is actually warranted versus when a PR is clean enough to leave alone.

The workflow’s configuration reflects that philosophy directly. It caps itself at 10 review comments and exactly one submitted review per run, scoped to COMMENT-level feedback only — it can flag concerns but can’t block a merge outright. It also shares a pr-review-base import with min-integrity: approved, meaning it won’t act on unverified or low-trust pull request content, and it pre-fetches diff data through a shared caching layer so repeated invocations on the same PR don’t re-download the same context.

Running on codex with the copilot/mai-code-1-flash-picker model, it’s tuned to be fast and cheap per invocation (roughly 3–8 AIC per run in these samples) rather than exhaustive — a reviewer that shows up quickly, says its piece if there’s something to say, and gets out of the way.

Complexity creep is invisible day-to-day and expensive in aggregate. A reviewer whose entire mandate is “is this more complicated than it needs to be?” is a narrow lens, but it’s one humans rarely apply consistently under review-fatigue. Ponytail Reviewer applies it on every ready-for-review PR, for free, every time.

Curious how it works under the hood? The workflow definition lives at .github/workflows/ponytail-reviewer.md in github/gh-aw. Explore the full catalog of agentic workflows, or build your own, at github.com/github/gh-aw.

Agent of the Day – September 1, 2026

Open pull requests have a way of quietly stalling — a CI check goes red and nobody notices, a review comment sits unanswered for a day, a branch drifts out of date. Today’s spotlight, PR Sous Chef, exists to catch exactly that kind of drift on gh-aw itself, checking in on every open, non-draft PR every fifteen minutes and nudging the Copilot coding agent only when there’s real work to hand back.

PR Sous Chef runs on the pi engine with openai/gpt-5.4, triggered on a tight every 15m schedule plus an on-demand /souschef slash command for anyone who wants to pull it into a specific PR conversation. It fetches all open PR branches (refs/pulls/open/*), reads through PR state, checks, and comments, and decides whether a targeted nudge is warranted — then posts a Copilot request if so.

Recent runs on September 1 show the pattern clearly: five runs in a single afternoon, all completed successfully, ranging from quiet passes with a single safe item to busier sweeps producing eight or nine actions each — see run #33509763563 and run #33516358980. Across its last five runs combined, it generated 20 safe-output items with zero errors and zero warnings — a workflow that does its job and gets out of the way.

The audit trail on that latest run is worth a closer look: 13 out of 13 automated quality graders passed clean, covering everything from tool-success-rate (100%) to loop detection (zero) to context-growth efficiency. The one flagged item was a handful of blocked outbound requests to github.com:443 — five out of fifty-three total network calls — a firewall-policy nuance rather than a functional problem, since the run still completed successfully and produced its full set of nudges.

What makes PR Sous Chef worth watching is its restraint. It doesn’t comment on every PR every cycle; a 15-minute schedule paired with a handful of safe items per run means most cycles are pure read-only reconnaissance — checking state, finding nothing actionable, and moving on. Only when a PR has genuinely gone quiet does it step in with a Copilot request, keeping the review queue moving without adding comment noise to PRs that are already progressing fine on their own.

It’s a small, unglamorous job — but in a repo with dozens of PRs in flight at any given time, having something check in every quarter-hour so nothing falls through the cracks is exactly the kind of quiet infrastructure that keeps a fast-moving project from stalling.


Curious how workflows like PR Sous Chef are built? Explore the project at github.com/github/gh-aw.