Blog

Sandbox Security Options Are Now Runtime Profiles

Sandbox security behavior is now selected through sandbox.agent.runtime. The separate sandbox.agent.legacy-security and sandbox.agent.sudo settings have been removed, making each runtime an explicit security and topology profile.

The default docker profile runs AWF without sudo and isolates network access. Workflows that need the previous privileged iptables behavior can select docker-sudo-iptables:

sandbox:
agent:
runtime: docker-sudo-iptables

This profile runs AWF with sudo, uses iptables-based networking, and permits host and GitHub Actions service access. It is required for sandbox.agent.allow-host-ports and for connecting to published services: ports. Other profiles retain their own isolation guarantees: gvisor adds kernel-level isolation, while docker-sbx and cloud-hypervisor use virtual-machine boundaries.

Run the fixer to update workflow frontmatter:

Terminal window
gh aw fix --write

The sandbox-runtime-profiles codemod rewrites this configuration:

sandbox:
agent:
legacy-security: enable

to:

sandbox:
agent:
runtime: docker-sudo-iptables

The codemod also removes obsolete sudo settings and preserves compatible runtime choices. When a combination cannot be migrated without changing its security intent, the fixer reports an actionable error instead of selecting a profile silently.

After migration, compile the workflow and review the generated lock file:

Terminal window
gh aw compile

See the sandbox configuration reference and agent runtime reference for the behavior and constraints of each profile.

Why the Built-In Playwright Tool Is Now CLI-Only

The built-in tools.playwright integration in gh-aw used to support two modes: a Docker-based MCP server (mode: mcp) and a CLI-based integration (mode: cli). As of this change, the built-in tool only supports CLI mode, and the compiler rejects mode: mcp with migration guidance instead of quietly starting a container. Workflows that still need the full Playwright MCP server can configure it explicitly under mcp-servers. Here is why CLI became the only built-in option.

MCP servers advertise their tools to the agent by loading a schema for every available function into the model’s context window. Playwright MCP exposes a wide surface of browser-automation tools — navigation, snapshots, clicks, evaluation, tracing, and more — and all of that schema has to be paid for in tokens on every turn, whether or not the agent uses most of it.

@playwright/cli instead exposes a single command, playwright-cli, that the agent invokes directly from bash with a subcommand such as goto, snapshot, or click. The agent only needs to have seen playwright-cli --help once (or installed skills via playwright-cli install --skills) to know how to drive the browser. There is no persistent tool schema competing with the rest of the workflow’s context for space, which matters for coding agents that also need room to reason about code, tests, and long-running tasks.

The built-in MCP mode ran Playwright inside a Docker container with its own image, arguments, and lifecycle that the compiler had to track, pin, and update independently. Every one of those knobs — container image version, extra MCP arguments, mounted volumes — was one more thing that could silently drift out of date or be misconfigured across workflows.

CLI mode collapses that surface. @playwright/cli is a single npm package installed directly on the runner, with one version to track and one command surface to allow through the shell permission system. Because it runs on the runner instead of in a separate container, it also reaches local development servers through localhost directly, without needing extra network plumbing between a container and the host. Less machinery means less to get wrong, and less to review when auditing what a workflow can do.

Some workflows genuinely benefit from MCP’s persistent state and richer introspection — for example, exploratory automation or self-healing tests that need to reason iteratively over page structure across many turns. That use case has not gone away; it is just no longer a hidden default. Configure it explicitly:

mcp-servers:
playwright:
command: npx
args:
- --yes
- "@playwright/mcp@0.0.79"
- --no-sandbox
allowed:
- browser_navigate
- browser_snapshot

Making this explicit means the dependency, the pinned version, and the exact allowed tool list are all visible in the workflow source, instead of being implied by a one-word mode: mcp setting. See the Playwright reference for the full migration table from MCP tool names to playwright-cli subcommands.

Agent of the Day – August 31, 2026

Every open-source repo eventually drowns in the same problem: a hundred open issues, half of them clearly related, none of them linked. Someone has to notice that six issues are all milestones of the same game plan, or that a dozen bug reports trace back to one root cause — and then actually go build the sub-issue tree. Today’s spotlight, Issue Arborist, is the daily workflow that does exactly that for gh-aw itself.

Every morning, Issue Arborist pulls the last 100 open issues that don’t already have a parent, reads through their titles, bodies, and labels, and decides what — if anything — deserves to be grouped. It’s deliberately conservative: it only proposes a brand-new parent issue when a cluster is strong and orphaned, and it only links a sub-issue when the relationship is unambiguous.

In its August 29 run, that discipline produced a genuinely useful cleanup. The agent noticed seven Deep Report issues — schema drift fixes, duplicate type consolidation, messaging tweaks — that were all clearly part of the same maintenance effort but had no tracking issue. Rather than link them to something that didn’t fit, it created a new parent, [Parent] Deep Report cleanup and schema consolidation, and attached all seven (#56706–#56712) underneath it. It then found two more clean clusters: four cache-miss issues that belonged under the existing Daily Cache Strategy Analyzer issue group (#56716), and six milestone issues for a fictional game project, “Neon Heist” (#56635), each one an obvious child of its own plan.

The very next day, on August 30, and again on August 31, it kept going — linking six more milestone issues for a different project, “Starling Drift” (#57144), and folding eight fresh [deep-report] findings into the existing issue-group parent (#56849), all without creating a single unnecessary new tracking issue.

What’s most interesting is what the agent chose not to do. Each run ends with a same-day discussion post explicitly listing the clusters it saw but declined to link — a group of near-duplicate “Implement skill-constraint-coverage” issues that might be duplicate attempts rather than a hierarchy, a set of repeated workflow-failure reports that could be separate incidents rather than one root cause, and several [aw] Smoke ... failed issues that looked plausible but weren’t confident enough to map automatically. That restraint is the whole point: a triage bot that links everything it can find isn’t useful, it’s noise. Issue Arborist’s value is in knowing when not to act.

Across its last three runs it made 129 link_sub_issue calls and opened one new parent issue — a quiet, steady gardening job that keeps gh-aw’s issue tracker legible without anyone having to ask it to.


Curious how workflows like Issue Arborist are built? Explore the project at github.com/github/gh-aw.

Weekly Update – August 31, 2026

Another packed week for github/gh-aw! We shipped three pre-releases (v0.87.5, v0.87.8, and v0.87.9) and merged well over a hundred pull requests spanning new workflow scheduling controls, compiler hardening, and a steady stream of trajectory grader implementations. Here’s what stood out.

The v0.87.9 line of pre-releases built on v0.87.8 and v0.87.5, focused on new workflow gating primitives, engine reliability, and internal grader tooling.

  • on.cooldown workflow gating (#56998): workflows can now declare a cooldown window so back-to-back triggers don’t pile up on the same target.
  • Typed on.stop-after field (#56983): stop-after now accepts GitHub Actions expressions in addition to static values, giving authors more flexible run-limiting logic.
  • Codex harness tool-schema diagnostics (#57256): unsupported-model tool-schema failures now surface with a dedicated, readable error message instead of a cryptic provider error.
  • Bump default MCP Gateway to v0.4.14 (#57188) and Bump Agentic Workflow Firewall to v0.28.10 (#56914): the usual steady drumbeat of dependency upgrades keeping the sandboxing and networking layers current.

Under the hood, the team also kept implementing new entries in the trajectory grader library — including event-entropy-rate, lempel-ziv-trajectory-complexity, policy-near-miss, exploration-error, and exploitation-error — building out a richer picture of how agents behave across runs.

AI Moderator is the quiet gatekeeper that watches newly opened issues, comments, and pull requests for spam, AI-generated noise, and link spam, then quietly labels or hides what it finds.

This week it stayed busy on the front lines — triggered repeatedly across incoming issues and PRs, including runs tied to the Codex harness fix (#57256) and a caveman instruction-verbosity pass. It runs read-only by design (no write-capable safe outputs get exercised unless it actually flags something), which is exactly the kind of low-risk, always-on moderation you want watching your front door.

Its recent runs did turn up as reliability “failures” in our observability logs — a reminder that even the calmest bouncer occasionally needs a coffee break, or in this case, a closer look at run classification before we assume the worst.

Usage tip: Because it runs read-only with threat-detection: false and tight per-window rate limits, ai-moderator is a solid template for any workflow that needs to watch high-volume public triggers (like issues: opened or pull_request: opened from forks) without risking runaway write actions.

→ View the workflow on GitHub

Update to v0.87.9 and give on.cooldown or the new intent-driven design guidance a try. As always, feedback and contributions are welcome in github/gh-aw.

Agent of the Day – August 28, 2026

Most teams write lint rules once and trust them forever. Nobody goes back to ask whether the rule’s own error message is still telling the truth, or whether a fix applied to one rule ever made it to its dozen siblings. Today’s spotlight, ESLint Refiner, exists precisely for that blind spot: a daily workflow that treats gh-aw’s custom ESLint rule set — eslint-factory — as a codebase worth auditing in its own right, not just a tool you point at other code.

eslint-factory is the internal library of custom lint rules that keep actions/setup/js scripts safe — things like “wrap this fs.mkdtempSync call in a try/catch” or “don’t compare objects with JSON.stringify equality.” Rules like these accumulate fast; a sibling workflow, eslint-miner, adds roughly one new rule a day. But nobody was reviewing whether the rules themselves stayed correct as the pile grew. ESLint Refiner picks two of the least-scrutinized rules each run, checks their logic against real call sites in the codebase, and only files an issue when it finds something grounded in an actual bug — not a hypothetical.

In its August 27 run, the agent reviewed require-mkdtempsync-try-catch and require-decodeuricomponent-try-catch, and turned up two real problems:

  1. A recurring overclaim. Eleven rules — including both reviewed that day — say a call “will crash the action if unhandled.” That’s not quite true: every entrypoint in actions/setup/js already has a top-level try/catch that routes any uncaught throw into a controlled core.setFailed, so nothing actually crashes silently. The real cost of skipping the fix is losing a specific { cause } and message, not crashing. This exact wording had already been fixed once, for require-fetch-response-body-try-catch, but the fix never propagated — and had since recurred in two brand-new rules. The agent filed issue #56288 asking for a reword-all pass plus a guard so it can’t quietly recur a third time.
  2. A misclassified literal. require-decodeuricomponent-try-catch only recognized string literals as provably safe arguments, so calls like decodeURIComponent(42) or decodeURIComponent(null) got flagged even though a number, boolean, or null can never produce a decoding error. Zero live call sites hit this today, but it’s a cheap, well-scoped fix worth closing before one does — filed as issue #56289.

Rather than silently move on, the agent also published a same-day discussion post summarizing exactly what it checked, what came back clean, and what’s queued for tomorrow’s review — including a note that its own repo-memory had gone stale for weeks even while it kept filing issues, which it then rebuilt from a ground-truth GitHub search.

The following day’s run, on August 28, kept the streak going with another clean pass — no errors, three safe outputs produced, business as usual for a workflow that’s quietly been doing this every day since it launched.

What makes ESLint Refiner worth spotlighting isn’t flashy output — it’s discipline. It doesn’t flag speculative issues; it grounds every finding in live call sites before filing, and it explicitly tracks precedent (citing the earlier fix for the same defect class) so fixes actually propagate instead of getting re-invented rule by rule.


Curious how workflows like ESLint Refiner are built? Explore the project at github.com/github/gh-aw.