Blog

Agent of the Day – September 22, 2026

Static analysis tools are supposed to be the trustworthy ones — the code that watches other code for mistakes. But a linter is still software, and software still has bugs, especially in the days right after it ships. Today’s Agent of the Day spends its nights making sure gh-aw’s own custom Go linters earn that trust: meet Sergo — the Serena Go Expert.

Sergo runs on a nightly schedule against the gh-aw repository, wired up to the Serena MCP language service for structural Go analysis rather than plain text search. Per its own execution plan, every run follows a disciplined loop: scan available Serena tools, load a bounded window of prior strategies from repo memory, split its effort 50/50 between reusing a proven approach and exploring something new, then generate up to three non-duplicate, evidence-backed issues before publishing a full report as a GitHub Discussion.

That discipline showed up clearly in its most recent run. Sergo noticed that uncheckedsliceindex — a Go linter merged just the day before via PR #62408 — had never been audited since landing. So it read the linter’s own source instead of trusting the green checkmark. The finding, filed as issue #62540, is precise: the linter’s core safety check, sameExpr, only recognizes a slice or string base when it reduces to a bare *ast.Ident. The moment the base is a struct field access like obj.Field[i] guarded by len(obj.Field), every one of the linter’s bounds-check recognizers — isLenOf, isInRangeLoop, isInBoundedForLoop — silently fails to connect the guard to the index, and a perfectly safe access gets flagged as unchecked.

Sergo didn’t stop at theory. It pointed to two real call sites already in the codebase — pkg/workflow/skills_ref_resolution.go:30-31 and pkg/workflow/cache_memory.go:355-356 — both ordinary, safely bounds-checked Go idioms that would trip this exact blind spot the moment the linter is wired into CI enforcement. It even checked the linter’s own test fixtures and confirmed they only ever index bare local variables, never a struct field, so the gap has zero test coverage today. The recommendation is equally concrete: generalize sameExpr to unwrap one level of SelectorExpr and compare the resolved field identity, rather than demanding a top-level identifier.

In the same run, Sergo also caught a recurrence: issue #62541 documents that a previously filed bug in bufioscannererunchecked — auto-closed as “not planned” by the repo’s issue-expiry workflow — was never actually fixed in code. Rather than silently letting the finding vanish, Sergo re-verified the source, confirmed the bug was still present line-for-line, and re-filed it with fresh evidence so it can’t quietly expire twice. Both issues were wrapped up with a daily discussion report summarizing the run’s strategy split, findings, and next-run focus — the kind of paper trail that turns one-off catches into a running audit history.

Zoom out across the last three nights and the pattern holds: three consecutive runs, three discussions, eight safe-output items total, zero errors. Sergo isn’t chasing style nits — it’s reading the linter’s own logic tree and asking “does this actually hold for code that isn’t in the test fixtures?” That’s a genuinely useful question to keep asking about static analysis tooling, especially the week after it ships.

Curious how a workflow like Sergo gets built on top of Serena’s language-service tools and gh-aw’s repo-memory system? Explore github/gh-aw and start writing your own agentic workflows today.

Weekly Update – September 21, 2026

Another busy week for github/gh-aw! We shipped a fresh release packed with reliability fixes, a smarter logs pipeline, and an updated model catalog — plus our usual dose of documentation polish.

v0.89.17 landed on September 19th, focused on hardening the AIC accounting pipeline, AWF/firewall integration, and safe-outputs handling.

  • Faster, smarter logs auditing — cached workflow runs are no longer redownloaded during logs audits (#61871), multi-target logs queries are now distributed fairly across targets (#61027), and per-run download duration/size is now tracked in an end-of-run stats summary (#60951).
  • Updated model catalog — added gemini-3.8-flash and claude-fable-5.1 aliases and corrected pricing for gpt-6-astra/gpt-5.6-sol (#61234).
  • MCP Gateway and firewall bumped — MCP Gateway updated to v0.4.25 (#61661) and gh-aw-firewall (AWF) updated to v0.28.20 (#61527) and v0.28.17 (#60945).
  • Better automatic grading — native Copilot tool calls are now included in the automatic grader trace payload for more accurate evaluation (#61426).
  • Fixed Code Scanning Fixer timeouts and tool denials (#61605).
  • Fixed slash command activation failing on CRLF line endings (#61602).
  • The safeoutputs CLI transport now fails loudly instead of silently failing open, surfacing real errors sooner (#61427).
  • Imported engine config (including auth) is now preserved when a workflow sets a top-level model (#61424).
  • Fixed several gaps in daily AIC (AI Credits) accounting, including legacy runs (#61313), pre-harness failures (#61232), and unassigned jobs (#61222).

Agent of the Week: deployment-incident-monitor

Section titled “ Agent of the Week: deployment-incident-monitor”

Meet deployment-incident-monitor, the on-call responder of the gh-aw fleet — it watches every deployment_status event and automatically files a deduplicated incident issue with root-cause analysis whenever something breaks.

This was its busiest week yet, firing 19 times — more runs than any other workflow in the repo. Most of the time it stayed quiet (no news is good news for deployments), but on September 19th it caught a real one: the Smoke Copilot - AOAI (Entra) workflow started failing with a 400 error because Azure OpenAI required organization verification for reasoning summaries. The agent didn’t just flag the failure — it traced it back to the exact commit, confirmed the change wasn’t the culprit, and filed issue #61892 with a full evidence trail linking the failing run and deployment.

Its incident history reads like a highlight reel of infra archaeology — Go toolchain mismatches, GitHub Pages queue timeouts, and cancelled sync_actions approvals — and thanks to close-older-issues, it keeps the incident tracker tidy instead of piling up duplicates.

Usage tip: Pair deployment_status-triggered monitors like this one with skip-if-match on your incident label so repeated failures from the same root cause collapse into a single, evolving issue instead of flooding your tracker.

→ View the workflow on GitHub

Update to v0.89.17 today, and check out the full changelog and prior releases on the releases page. As always, feedback and contributions are welcome in github/gh-aw.

Agent of the Day – September 15, 2026

Custom lint rules are a strange kind of code: they’re supposed to catch other people’s mistakes, but nobody is watching them for their own. A rule ships, looks reasonable in review, and then quietly misfires on some edge case nobody tested — a false positive that either gets ignored forever or trains developers to distrust the whole linter. Today’s Agent of the Day exists specifically to close that gap: ESLint Refiner, gh-aw’s daily auditor of its own custom rules.

ESLint Refiner runs every morning against eslint-factory, the TypeScript package that defines gh-aw’s custom ESLint rules for the JavaScript helpers under actions/setup/js/. Its job, per its own workflow definition, is narrow and disciplined: review recent diagnostics, spot false positives or weak checks, propose up to three concrete refinement tasks, file non-duplicate issues with acceptance criteria, and publish a daily discussion summarizing what it found — all while persisting its findings to repo memory so tomorrow’s run doesn’t repeat today’s work.

In its most recent run, that discipline paid off with a genuinely subtle catch. The rule prefer-actions-exec-over-child-process flags any child_process.exec/execSync/execFile call inside a file that carries the <reference types="@actions/github-script" /> marker, on the assumption that such files always run inside a GitHub Actions github-script step where the @actions/exec global is available. ESLint Refiner traced two files — merge_remote_agent_github_folder.cjs and build_checkout_manifest.cjs — where that assumption breaks down: both also require("./shim.cjs"), a helper whose own docstring says it exists precisely so github-script-flavored modules can run standalone inside the safe-outputs and MCP-scripts servers. Crucially, shim.cjs polyfills core and context, but not exec — so the rule’s suggested rewrite to @actions/exec isn’t even available in that execution mode, making the diagnostic a false positive.

Rather than filing a vague “please double-check this rule” ticket, the agent produced issue #61044 with exact line numbers (194, 198, 207, 212, 215 in one file; 52 and 61 in the other), a contrasting true-positive example (get_current_branch.cjs, which carries the same marker but has no shim.cjs fallback and is correctly flagged), a proposed requiresShimCjs() guard for the rule’s source, and acceptance criteria that explicitly protect against over-correcting into a blanket suppression. It then wrapped the run with a daily discussion report summarizing the finding for anyone tracking the linter’s health over time.

The run wasn’t flawless — gh-aw’s own audit tooling flagged it as resource-heavy for its task domain, burning 83 turns and 89k tokens against a 45-minute budget, with one blocked network request to api.anthropic.com. That’s useful signal in its own right: even a workflow that ships a precise, well-evidenced result can still be a candidate for tightening, and gh-aw’s audit pipeline calls that out automatically rather than letting a “success” status hide the cost.

That combination — a narrow domain, line-level evidence, a fix proposal that anticipates its own failure modes, and a daily paper trail in Discussions — is exactly the kind of quiet, compounding value an “agent of the day” should demonstrate. Rules that lint your code deserve a linter of their own, and ESLint Refiner is gh-aw’s answer to that.

Want to see how workflows like ESLint Refiner are built? Check out github/gh-aw and start writing your own agentic workflows today.

Agent of the Day – September 14, 2026

Every AI-heavy codebase has the same quiet failure mode: a provider ships a new flagship model, nobody updates the alias map, and workflows silently fall back to a stale default until someone notices the bill or the benchmark looks off. Today’s Agent of the Day exists specifically to close that gap before it opens — the Daily Model Inventory Checker .

Agent of the Day: Daily Model Inventory Checker

Section titled “Agent of the Day: Daily Model Inventory Checker”

This scheduled gh-aw workflow (.github/workflows/daily-model-inventory.md) runs once a day and treats model tracking like a real audit, not a scrape-and-hope job. It pulls live catalogs from OpenAI, Anthropic, and Google directly via their APIs, uses the built-in AWF /reflect endpoint to query Copilot’s own model metadata, and cross-references everything against pkg/workflow/data/model_aliases.json and pkg/cli/data/models.json — the two files that decide which model a workflow actually gets when it asks for large or agent.

Its last three scheduled runs — Sept 11, Sept 12, and Sept 13 — all completed cleanly, with zero errors and zero warnings across the board, spending 13 to 18 minutes and 35k–75k tokens per run to reconcile roughly 240 models across five providers each time.

The workflow’s real value shows up when it finds something. On September 9, it surfaced issue #59709, “Model alias inventory update - 2026-09-09,” reporting that OpenAI’s new flagship gpt-6-astra was live in both the Copilot and OpenAI catalogs but wasn’t covered by any existing alias glob — meaning workflows requesting large or agent would silently miss the newest, most capable model. The issue laid out exact pricing ($1,000 / $5,000 per million tokens, input/output), the precise alias diff needed, and a full breakdown of all 242 models it had checked across every provider, including which legacy IDs to deliberately leave alone.

That report turned directly into two merged pull requests: #59711, which added the gpt-6 alias pattern and wired gpt-6-astra into the large and agent resolution chains along with its pricing metadata, and #59703, a same-day fix bumping the Copilot SDK dependency after a version mismatch had briefly broken the inventory job itself. Both merged within roughly 15 minutes of being opened.

What stands out about this agent isn’t volume — it’s precision and restraint. Its own report explicitly calls out that no other alias gaps existed that day, walks through why each existing wildcard (gemini-*flash*, gpt-5.x, sonnet/opus/haiku) still correctly catches every other new model observed, and preserves historical entries like claude-opus-4.5 and gpt-4.1 rather than pruning them, in case older workflows still reference them. It also cross-validates pricing from two independent sources — Copilot SDK billing data and the reflect endpoint — and only proposes a change when both agree. When nothing needs fixing, it stays quiet; when something does, it hands maintainers a fully-sourced diff instead of a vague nudge.

For a project that runs hundreds of AI-driven workflows depending on consistent model resolution, that kind of nightly discipline is the difference between a broken default nobody notices for weeks and a same-day, two-PR fix.

Agentic workflows activity

Curious how workflows like this one are built? Check out github/gh-aw and see what your next agent could catch.

Weekly Update – September 14, 2026

It was a fast-moving week for github/gh-aw! The team shipped seventeen releases — from v0.88.5 all the way to v0.89.12 — packed with gh aw logs observability upgrades, model-routing fixes, and a steady drumbeat of CI and credential-security hardening.

v0.89.0 was the headline release of the week, focused on hardening agentic engine model selection, improving gh aw logs observability, and tightening safe-output guardrails.

  • Identifiable MCP tool calls in logs (#59579): gh aw logs --json now records the timestamp, server name, and tool name for every MCP call, making it far easier to trace which server and tool produced a given usage entry.
  • --ignore-workflow-runs for gh aw logs (#59697): exclude specific runs (by numeric ID or slug/ID) from log collection without shrinking your requested result count.
  • Refreshed cached logs JSON (#59690): gh aw logs --cached-json now replaces the cache file with up-to-date results after each successful collection instead of leaving it stale.
  • GPT-6 Astra model support (#59711): gpt-6-astra is now recognized in model alias resolution and pricing catalogs for GitHub Copilot and OpenAI.
  • Fixed threat detection reporting config_error for workflows using custom engines (#59636), so custom-engine workflows now get proper threat analysis instead of silently skipping it.
  • Fixed a Copilot SDK model inventory collection break caused by an incompatible CLI platform package after a dependency bump (#59703).
  • Bundled MCP gateway upgraded to v0.4.20 (#59602), including a safe-outputs sink-visibility exemption fix.

The week closed out with a small but important security release, v0.89.12:

  • Reduced credential blast radius in the slash-command router (#60685): the generated central slash-command router workflow now checks out the repository with persist-credentials: false, so GITHUB_TOKEN is no longer persisted in local git config for the lifetime of the routing job.
  • Fixed the “Integration: CMD Tests” CI job (#60683), restoring a green CI signal.

Between the two headline releases, dozens of PRs kept the fleet humming:

Meet cli-version-checker — the fleet’s diligent version scout, running daily to watch for new releases of Claude Code, GitHub Copilot CLI, OpenAI Codex, the GitHub MCP Server, Playwright CLI, MCP Gateway, Pi, threat-detect, and a stack of container-scanning tools like actionlint, syft, grype, and zizmor.

Over its last three scheduled runs this agent had a genuinely mixed week: one clean 6.6-minute pass that filed its usual “[ca]“-prefixed update issue, one run that failed after just 49 seconds, and one earlier run that took 5.6 minutes before also hitting trouble. All told it burned through roughly 40K tokens and made 26 GitHub API calls chasing down version numbers across nine different tools and eight container images — a lot of bookkeeping for one little agent.

Its self-imposed 2-day issue expiry is a nice touch: if nobody acts on a version bump quickly, the checker doesn’t let stale “you should upgrade” nags pile up in the issue tracker forever.

Usage tip: For any “check external state and file an issue” workflow, pair a short expires window on the safe-output with a cookie label — it keeps the backlog honest and makes triage-by-label trivial.

→ View the workflow on GitHub

Check out the latest releases of gh-aw and give the new gh aw logs observability features a spin. As always, feedback and contributions are welcome in github/gh-aw.