Artifacts

GitHub Agentic Workflows upload several artifacts during workflow execution. This reference documents every artifact name, its contents, and how to access the data — especially for downstream workflows that use gh run download directly instead of gh aw logs.

Artifact NameConstantTypeDescription
agentconstants.AgentArtifactName
Source: pkg/constants/job_constants.go
Multi-fileUnified agent job outputs (logs, safe outputs, token usage summary)
agent-output-fallbackconstants.AgentOutputFallbackArtifactNameMulti-fileSmall dedicated copy of the processed agent output (agent_output.json) and raw safe-output NDJSON (safeoutputs.jsonl), used when the larger agent upload fails or times out
activationconstants.ActivationArtifactNameMulti-fileActivation job output (aw_info.json, prompt.txt, rate limits)
firewall-audit-logsconstants.FirewallAuditArtifactName
Source: pkg/constants/constants.go
Multi-fileAWF firewall audit/observability logs (token usage, network policy, audit trail)
detectionconstants.DetectionArtifactNameConditionalLegacy inline engine (features.gh-aw-detection: false): single-file detection.log. The default external gh-aw-detection engine: multi-file detection_result.json + step-summary.md; detection.log is intentionally not uploaded (see below)
safe-outputconstants.SafeOutputArtifactNameLegacy/back-compatHistorical standalone safe output artifact (safe_output.jsonl); in current compiled workflows this content is included in the unified agent artifact instead
agent-outputconstants.AgentOutputArtifactNameLegacy/back-compatHistorical standalone agent output artifact (agent_output.json); in current compiled workflows this content is included in the unified agent artifact instead
infoconstants.InfoArtifactNameSingle-fileStandalone copy of the workflow run information (aw_info.json), uploaded by the activation job in addition to the copy bundled in the activation artifact
aw-info—Legacy/back-compatHistorical standalone engine-configuration artifact (aw_info.json); current compiled workflows upload this as info instead
prompt—Legacy/back-compatHistorical standalone prompt artifact (prompt.txt); in current compiled workflows prompt.txt is included in the activation artifact instead
experimentconstants.ExperimentArtifactNameMulti-fileA/B experiment state (state.json) uploaded by the activation job when experiments are declared in the frontmatter
usageconstants.UsageArtifactNameMulti-fileCompact conclusion-job artifact with workflow-run metadata and token-usage files used by lightweight reporting and forecasting paths
evalsconstants.EvalsArtifactNameSingle-fileBinEval evaluation results (evals.jsonl) uploaded by the evals job when evals are declared in the workflow frontmatter
safe-outputs-itemsconstants.SafeOutputItemsArtifactNameMulti-fileSafe output items manifest (safe-output-items.jsonl), temporary ID map (temporary-id-map.json), and failure diagnostics (safe-output-errors.json, written only when the Process Safe Outputs step fails)
code-scanning-sarifconstants.SarifArtifactNameSingle-fileSARIF file for code scanning results

The gh aw logs and gh aw audit commands support --artifacts to download only specific artifact groups:

Set NameArtifacts DownloadedUse Case
allEverythingFull analysis (default)
agentagentAgent logs and outputs
activationactivationActivation data (aw_info.json, prompt.txt)
firewallfirewall-audit-logsNetwork policy and firewall audit data
mcpfirewall-audit-logsMCP gateway traffic logs
detectiondetectionThreat detection output
experimentexperiment, usageA/B experiment state (only present when experiments are declared)
usageusageCompact conclusion-job artifact for lightweight reporting and forecasting
evalsusageBinEval evaluation results (only present when evals are declared)
gradersusage, agent, agent-output-fallbackDeterministic grader results (only present when graders are declared)
github-apiactivation, agentGitHub API rate limit logs
Terminal window
# Download only firewall artifacts
gh aw logs <run-id> --artifacts firewall
# Download agent and firewall artifacts
gh aw logs <run-id> --artifacts agent --artifacts firewall
# Download everything (default)
gh aw logs <run-id>

The firewall-audit-logs artifact is uploaded by all firewall-enabled workflows. It contains AWF (Agent Workflow Firewall) structured audit and observability logs.

! Important: This artifact is separate from the agent artifact. Token usage data (token-usage.jsonl) lives here, not in the agent artifact.

firewall-audit-logs/
├── api-proxy-logs/
│ ├── token-usage.jsonl ← Token usage data (input/output/cache tokens per API request)
│ └── token-diag.log ← Token diagnostics JSONL (only when AWF_DEBUG_TOKENS=1)
├── squid-logs/
│ └── access.log ← Network policy log (domain allow/deny decisions)
├── audit.jsonl ← Firewall audit trail (policy matches, rule evaluations)
└── policy-manifest.json ← Policy configuration snapshot

token-diag.log is written by the AWF api-proxy diag() path (containers/api-proxy/token-persistence.js) to $AWF_TOKEN_LOG_DIR/token-diag.log (default /var/log/api-proxy/token-diag.log). It is only emitted when AWF_DEBUG_TOKENS=1, so set that environment variable on the workflow step that runs with AWF enabled when you need token diagnostics.

Recommended: Use gh aw logs

Terminal window
# Download and analyze firewall data
gh aw logs <run-id> --artifacts firewall
# Output as JSON for scripting
gh aw logs <run-id> --artifacts firewall --json

Direct download with gh run download:

Terminal window
# Download the firewall-audit-logs artifact
gh run download <run-id> -n firewall-audit-logs
# Token usage data is at:
cat firewall-audit-logs/api-proxy-logs/token-usage.jsonl
# Network access log is at:
cat firewall-audit-logs/squid-logs/access.log
# Audit trail is at:
cat firewall-audit-logs/audit.jsonl
# Policy manifest is at:
cat firewall-audit-logs/policy-manifest.json

Downstream workflows sometimes download agent-artifacts or agent expecting to find token-usage.jsonl. This will silently return no data — the token usage file is only in the firewall-audit-logs artifact.

Terminal window
# ✗ WRONG — token-usage.jsonl is NOT in the agent artifact
gh run download <run-id> -n agent
cat agent/token-usage.jsonl # File not found!
# ✓ CORRECT — download from firewall-audit-logs
gh run download <run-id> -n firewall-audit-logs
cat firewall-audit-logs/api-proxy-logs/token-usage.jsonl

The JSONL files in this artifact are described by versioned JSON Schemas published by github/gh-aw-firewall. Each record includes a _schema field (for example "audit/v0.26.0") so consumers can identify the record type and AWF version.

FileSchema assetPinned URL
audit.jsonlaudit.schema.jsonhttps://github.com/github/gh-aw-firewall/releases/download/<tag>/audit.schema.json
api-proxy-logs/token-usage.jsonltoken-usage.schema.jsonhttps://github.com/github/gh-aw-firewall/releases/download/<tag>/token-usage.schema.json

Use releases/latest/download/ in place of a specific tag to track the most recent published release. Schemas are versioned by AWF release tag; consumers should match _schema by prefix (for example _schema.startsWith("audit/")) so additive changes remain non-breaking.

The unified agent artifact contains agent job outputs:

  • Agent execution logs
  • Safe output data (agent_output.json)
  • GitHub API rate limit logs (github_rate_limits.jsonl)
  • Token usage summary (agent_usage.json) — aggregated totals only; per-request data is in firewall-audit-logs. When AWF records include valid ai_credits_this_response and ai_credits_total values, the summary preserves those reported values instead of repricing the tokens.
  • otel.jsonl — OTLP span mirror written by gh-aw’s JavaScript span exporters when observability.otlp is configured

For OTLP configuration, runtime environment variables, and span semantics, see the OpenTelemetry guide.

The activation artifact contains activation job outputs:

  • aw_info.json — Engine configuration and workflow metadata
  • prompt.txt — The generated prompt sent to the AI agent
  • github_rate_limits.jsonl — Rate limit data from the activation job

The detection artifact is conditional:

  • Inline engine (default): detection.log, the threat-detection analysis output. Legacy name: threat-detection.log.
  • External gh-aw-detection engine (the default, or features.gh-aw-detection: true): detection_result.json and step-summary.md.

The experiment artifact is uploaded by the activation job only when the workflow frontmatter declares one or more experiments entries. It contains:

  • state.json — Cumulative per-variant invocation counters used to balance A/B assignments across runs

The conclusion job also copies the experiment state (state.jsonl or state.json) and the current run’s assignments.json into the experiment/ directory of the usage artifact, so gh aw audit can report experiment assignments from the usage artifact alone.

Terminal window
# Download the experiment artifact for a specific run
gh aw audit <run-id> --artifacts experiment
# Display the A/B experiment section in the audit report
gh aw audit <run-id>

The A/B Experiments section of the audit report shows the variant chosen for the run and the cumulative counts:

A/B Experiments
• style = concise (cumulative: concise:5, detailed:4)

See A/B Experiments for how to declare experiments in workflow frontmatter.

The usage artifact is a compact conclusion-job artifact with workflow-run metadata and token-usage files for lightweight reporting and forecasting, so downstream tools can read aggregated usage data without downloading the full agent artifact.

Its activity/summary.json file uses the usage-activity-summary/v1 schema. The optional activity sections are additive; the working_set and friction sections are always written when the calculation step executes. Runs produced before a section shipped simply omit it, and every consumer treats a missing section as unmeasured:

The ledger.transactions_added count covers repo-memory ledger appends recorded as ledger_mutation items in the downloaded safe-outputs manifest; queued ledger_append safe outputs are not counted because they are persisted later in a separate job. It is zero when that manifest is present without ledger mutations and absent when the manifest is unavailable. Ledger compaction runs in Agentic Maintenance, not in agent workflow runs; its plan and apply results appear in the maintenance job summaries.

{
"schema": "usage-activity-summary/v1",
"ledger": {
"transactions_added": 3
},
"firewall": {
"total_requests": 12,
"allowed_requests": 10,
"blocked_requests": 2
},
"gateway": {
"total_calls": 5,
"failed_calls": 1,
"total_input_size": 1000,
"total_output_size": 5000,
"max_input_size": 400,
"max_output_size": 3000,
"tool_calls": [
{
"tool_call_id": "call-1",
"timestamp": "2026-09-09T00:00:00Z",
"server_name": "github",
"tool_name": "issue_read",
"request_size": 200,
"response_size": 800,
"duration_ms": 100,
"outcome": "success"
}
],
"servers": [
{
"server_name": "github",
"request_count": 5,
"tool_call_count": 5,
"failed_calls": 1
}
],
"tools": [
{
"server_name": "github",
"tool_name": "issue_read",
"call_count": 5,
"failed_calls": 1,
"total_input_size": 1000,
"total_output_size": 5000,
"max_input_size": 400,
"max_output_size": 3000,
"avg_duration_ms": 120,
"max_duration_ms": 250
}
]
},
"integrity": {
"total_filtered": 2,
"filtered_server_counts": { "github": 2 },
"filtered_tool_counts": { "issue_read": 2 },
"filtered_reason_counts": { "integrity": 2 }
},
"steering": {
"total_events": 3,
"event_counts": {
"token_steering": 2,
"timeout_steering": 1
}
},
"working_set": {
"measurement_state": "measured",
"rebuild_factor": 3.9017857142857144,
"cumulative_input_tokens": 874000,
"peak_input_tokens": 224000,
"rebuild_excess_tokens": 650000,
"invocations": 5
},
"friction": {
"measurement_state": "statistical",
"canonical_unit": "aic",
"sources": ["agent_session", "agent_token_usage", "firewall", "mcp_gateway"],
"total_events": 2,
"total_occurrences": 4,
"counted_occurrences": 3,
"suppressed_occurrences": 1,
"linked_invocations": 1,
"unattributed_occurrences": 0,
"cost": {
"aic": 1.25,
"tokens": { "input": 200, "output": 40, "cache_read": 10, "cache_write": 2, "reasoning": 4, "total": 256 },
"turns": 1,
"tool_calls": 3,
"latency_ms": 550
},
"dimension_states": {
"aic": "statistical",
"tokens": "statistical",
"turns": "measured",
"tool_calls": "measured",
"latency_ms": "measured"
},
"uncertainty": {
"aic": {
"state": "statistical",
"method": "mean_invocation_apportionment",
"confidence": "low",
"relative_error": 0.3,
"sample_size": 2,
"basis": "AI credits of errored invocations (measured), of linked follow-up invocations (causal), or the mean healthy-invocation cost (statistical)"
}
},
"drivers": [
{
"driver": "mcp_tool_error",
"class": "tool_failure",
"source": "mcp_gateway",
"events": 1,
"occurrences": 1,
"counted_occurrences": 1,
"suppressed_occurrences": 0,
"state": "causal",
"cost": { "aic": 0.5, "tokens": { "input": 100, "output": 28, "cache_read": 0, "cache_write": 0, "reasoning": 0, "total": 128 }, "turns": 0, "tool_calls": 1, "latency_ms": 250 }
}
],
"groups": [
{
"group_id": "tool_failure",
"primary_source": "mcp_gateway",
"event_ids": ["mcp_tool_error:call-1"],
"total_occurrences": 4,
"counted_occurrences": 3,
"suppressed_occurrences": 1,
"rule": "highest-fidelity source owns overlapping occurrences; lower-fidelity sources contribute only their excess"
}
],
"events": [
{
"id": "mcp_tool_error:call-1",
"driver": "mcp_tool_error",
"source": "mcp_gateway",
"group_id": "tool_failure",
"label": "github/issue_read",
"timestamp": "2026-09-09T00:00:01Z",
"occurrences": 1,
"counted_occurrences": 1,
"suppressed_occurrences": 0,
"state": "causal",
"dimension_states": {
"aic": "causal",
"tokens": "causal",
"turns": "unsupported",
"tool_calls": "measured",
"latency_ms": "measured"
},
"cost": { "aic": 0.5, "tokens": { "input": 100, "output": 28, "cache_read": 0, "cache_write": 0, "reasoning": 0, "total": 128 }, "turns": 0, "tool_calls": 1, "latency_ms": 250 }
}
],
"unmeasured_drivers": [{ "driver": "firewall_block", "reason": "no_occurrences" }]
},
"safe_outputs": {
"total_items": 2,
"items_by_type": {
"create_issue": 1,
"add_labels": 1
},
"items": [
{
"type": "create_issue",
"provider": "github",
"url": "https://github.com/owner/repo/issues/42",
"number": 42,
"repo": "owner/repo",
"target": {
"provider": "github",
"repository": "owner/repo",
"number": 42
},
"timestamp": "2026-09-14T00:00:00.000Z"
},
{
"type": "add_labels",
"provider": "github",
"number": 42,
"repo": "owner/repo",
"target": {
"provider": "github",
"repository": "owner/repo",
"number": 42,
"kind": "issue"
},
"labels": [
{
"name": "triage",
"database_id": 1234,
"node_id": "LA_example"
}
],
"timestamp": "2026-09-14T00:00:01.000Z"
}
]
}
}

MCP tool_calls contain quantitative metadata only. Tool-call IDs are replaced with run-local opaque identifiers, and request and response content is never copied into the usage artifact. outcome is success, failure, or incomplete.

safe_outputs.items contains the provider-neutral records from the safe-output manifest. Each record identifies the provider and operation and includes the available URL, repository, number, provider ID, human-readable identifier, target, and label details. This lets consumers reconstruct created GitHub, Jira, Linear, and other provider entities from the usage artifact without querying those services. Label records include the associated issue or pull request, label name, and GitHub database or node ID when returned by the API.

The conclusion job derives gateway and integrity from MCP gateway logs, falling back to rpc-messages.jsonl when gateway.jsonl is unavailable. These compact aggregates let gh aw logs --artifacts usage report MCP call, payload-size, duration, failure, and integrity-filter metrics without downloading raw logs. Cross-run reports include runs_with_filtered_events; the existing logs report summary remains the source for the total number of runs.

The steering section aggregates AWF API proxy events by normalized event name. gh aw audit --artifacts usage exposes these counters in firewall_token_usage.steering_event_counts, so steering behavior can be inspected without downloading raw firewall logs.

rebuild_factor is cumulative_input_tokens / peak_input_tokens, where each invocation contributes the canonical input_tokens value from the agent token_usage.jsonl record. Cache-read and cache-write fields are not added because provider normalization has already produced that logical input count. The factor is omitted when measurement_state is unavailable; partial means usable records were measured but malformed or unsupported records were ignored.

Working-Set Rebuild Factor measures cumulative context reconstruction relative to peak invocation context. It is an efficiency/trajectory metric, not a measurement of semantic coherence debt and not a predictor of task success. It cannot identify missing task facts or classify outcome quality. The metric is conceptually inspired by “The Working Set of a Coding Agent: Coherence Debt in Repository-Scale Tasks”, while deliberately limiting the implementation to observable token traffic.

Friction cost is the estimated avoidable marginal cost attributable to an execution-friction event, relative to the counterfactual execution in which that event did not occur.

The friction section contains the precomputed friction cost for tool calls that failed, responses that were filtered, requests that were blocked, and model invocations that errored or were retried. The conclusion job computes it once from the logs it already parses, so consumers that download only the usage artifact read the finished numbers instead of re-deriving them. gh aw logs and gh aw audit always prefer this precomputed section when it is present.

AI credits (aic) are the canonical unit. Every other dimension — token classes, turns, tool calls, and latency — is reported only for drivers that support it, and unsupported is stated explicitly rather than reported as zero.

Statistical token estimates retain fractional precision across attributed occurrences, then emit integer token counts with deterministic remainder distribution across event records.

StateMeaning
measuredThe cost was read directly off the record describing the friction event.
causalThe cost was taken from the model invocation the event demonstrably caused (the first unconsumed invocation after the event).
statisticalThe cost was apportioned from the mean healthy-invocation cost of the same run.
unavailableThe driver was detected, but no cost could be attributed with the data available.
unsupportedThe driver cannot express this dimension at all.

measurement_state at the top of the section is the state of the canonical aic dimension and equals the weakest state contributing to it. A run with friction sources but no friction reports measured with a zero cost; a run with no usable source reports unavailable.

Consequently, a counted event with unavailable AIC makes the run-level state unavailable even when other events have measured or estimated costs. The numeric cost remains the sum of attributed portions and does not imply that the unattributed portion is zero; unattributed_occurrences records occurrences whose AIC could not be attributed.

DriverClassSourceAICTokensTurnsTool callsLatency
mcp_tool_errortool_failuremcp_gatewayderivedderivedunsupportedmeasuredmeasured
session_tool_failuretool_failureagent_sessionderivedderivedunsupportedmeasuredunsupported
integrity_filterintegrity_filtermcp_gatewayderivedderivedunsupportedmeasuredunsupported
firewall_blocknetwork_blockfirewallderivedderivedunsupportedunsupportedunsupported
agent_api_errormodel_erroragent_token_usagemeasuredmeasuredmeasuredunsupportedmeasured

measured means the dimension is read from the friction record itself. derived means the dimension is attributed causally when a follow-up invocation can be linked, and statistically otherwise. unsupported means no data source expresses that dimension for the driver. Firewall-block costs have no invocation-level causal link, but AIC and tokens are statistically estimated from healthy invocations when that baseline is available; otherwise those dimensions are unavailable. Drivers whose source is absent, or that produced no occurrences, are listed in unmeasured_drivers with a no_occurrences or source_unavailable:<source> reason.

When total_run_aic_partial is true, total_run_aic sums only invocations with AIC telemetry, and friction_ratio divides attributed friction AIC by that partial denominator. Because the numerator can include causal or statistical estimates, this ratio can exceed 1 and must not be interpreted as a bounded share of full run cost. Malformed token-usage records make the run total and ratio unavailable rather than publishing a misleading partial denominator.

Events are grouped by causal class (group_id) so that the same underlying failure observed by several sources is counted once. Within a group, sources are ranked by fidelity — agent_token_usage, then mcp_gateway, then agent_session, then firewall — and the highest-fidelity source owns the overlapping occurrences. Lower-fidelity sources contribute only the occurrences they observed in excess, reported as counted_occurrences with the remainder in suppressed_occurrences and suppressed_by. A single model invocation is likewise linked to at most one friction event across the whole run, so causal attribution can never bill the same AI credits twice. Grouping depends only on the observed counts and a fixed fidelity order, so the same input always produces the same output.

The event_ids list references only entries included in the capped events array. A group with omitted IDs sets event_ids_truncated to true; its occurrence and cost aggregates still include every event.

Each dimension carries an uncertainty entry with its state, the method used (direct_record, next_invocation_linkage, mean_invocation_apportionment, or none), a confidence bucket, the sample_size behind the estimate, and bounds where supported. Measured bounds equal the observed value. Causal estimates range from zero to the linked invocation cost because the exact counterfactual is not observable. Statistical estimates use a 95% interval based on the relative standard error of healthy invocations; bounds are omitted when fewer than two healthy invocations were available.

Friction cost is an efficiency signal, not an attribution of blame, a claim that the underlying action was unnecessary, or a prediction of task success. Statistical attribution assumes the cost of recovering from friction resembles the average invocation of the same run, which is an approximation, not a measurement.

The usage artifact also carries experiment and evals data when the workflow declares them, so gh aw audit --artifacts usage can mine both without downloading other artifacts:

  • experiment/state.jsonl, experiment/state.json, experiment/assignments.json — A/B experiment state and the current run’s variant assignments
  • evals.jsonl, evals/token_usage.jsonl, evals/execution.json — BinEval results, evals token usage, and evals execution evidence
  • detection/detection_result.json — When threat detection is enabled, the detection job result, conclusion, categorized failure reason, and validated threat verdict flags (when available). Raw detector reasons and logs are not copied into this file. gh aw audit --artifacts usage reports failed or warned detection and detected threats as security findings.

Token-usage files are diagnostic data produced in the agent runtime. Their mirrored AIC fields support usage reporting and analysis, but are not sufficient evidence to classify a provider failure as a trusted budget-enforcement event.

Terminal window
# Download only the usage artifact
gh aw logs <run-id> --artifacts usage
# Or with gh run download
gh run download <run-id> -n usage

The evals artifact is uploaded by the evals job only when the workflow frontmatter declares one or more evals entries. It is not present on runs without evals and contains:

  • evals.jsonl — Per-question BinEval evaluation results (YES/NO records) produced by running the declared evaluation questions against the agent output
Terminal window
# Download only the evals artifact
gh aw logs <run-id> --artifacts evals
# Or with gh run download
gh run download <run-id> -n evals

The gh aw audit command exposes an --evals flag that skips runs without evals results and automatically downloads the evals artifact when --artifacts is narrowed:

Terminal window
# Audit only runs that contain evals results
gh aw audit <run-id> --evals

When the workflow frontmatter declares one or more graders, the grader files are stored inside existing artifacts rather than uploaded as a standalone graders artifact. They live under agent/graders/ in the unified agent artifact, under graders/ in the agent-output-fallback artifact when the fallback transport is used, and under usage/graders/ after the conclusion job mirrors them for lightweight downloads. The files are:

  • grader_manifest.json — The configured graders with their unit, direction, and threshold
  • grader_results.json — The validated grader results (status, raw value, pass/fail) computed from the run trace
Terminal window
# Download the artifacts that carry grader results
gh aw logs <run-id> --artifacts graders
# Or with gh run download against the actual artifacts
gh run download <run-id> -n usage
gh run download <run-id> -n agent
gh run download <run-id> -n agent-output-fallback

gh aw audit reports grader outcomes in its console output and includes them in the JSON report under the graders key (results, plus total, passed, failed, error_count, and unavailable_count). The artifact also retains the exact operational-value evaluator bytes used for that run.

Artifact names changed between upload-artifact v4 and v5. The gh aw logs and gh aw audit commands handle both naming schemes transparently:

Old Name (pre-v5)New Name (v5+)File Inside
aw_info.jsonaw-infoaw_info.json
safe_output.jsonlsafe-outputsafe_output.jsonl
agent_output.jsonagent-outputagent_output.json
prompt.txtpromptprompt.txt
threat-detection.logdetectiondetection.log (inline engine only)

Single-file artifacts are automatically flattened to root level regardless of their artifact directory name. Multi-file artifacts (firewall-audit-logs, agent, activation, experiment, and detection when the external gh-aw-detection engine is enabled) retain their directory structure.

When workflows are invoked via workflow_call, GitHub Actions prepends a short hash to artifact names (e.g., abc123-firewall-audit-logs). The CLI handles this automatically by matching artifact names that end with -{base-name}.

Terminal window
# Both of these are recognized as the firewall artifact:
# - firewall-audit-logs (direct invocation)
# - abc123-firewall-audit-logs (workflow_call invocation)

See Audit Commands for downloading and analyzing workflow run artifacts, Cost Management for token-usage and spend reporting, Network for firewall configuration, and Compilation Process for how workflows upload artifacts.