# Metergrade Setup Kit

This kit checks the economics of your Azure API Management AI traffic. It
is plain text, KQL, and a workbook definition — read it before running it.

## What this is not

No installer. No daemon. No agent runtime. No outbound network calls to
Metergrade. No credentials.

## Two ways to use it

**Economic Control Check** (no account, nothing transferred) — import the
workbook, run the queries, read the results in your own Azure environment.

**Metergrade Free** (saved baseline) — produce a supported export locally,
sign in to Metergrade, and upload it yourself.

## Before you start: what has to be switched on

The check reads two Azure Monitor tables, and both are produced by APIM
diagnostic settings pointed at a Log Analytics workspace:

| Table | Carries | Diagnostic category |
|---|---|---|
| `ApiManagementGatewayLogs` | request identity, operation, status, latency | Gateway logs |
| `ApiManagementGatewayLlmLog` | model, deployment, token counts | **LLM logs — separate, and off by default** |

**The LLM category is the one that is usually missing.** It is enabled
separately from gateway logging, and a workspace that has never received it
does not contain an empty table — it contains no table at all, so a query
against it fails to resolve rather than returning zero rows.

That is a configuration fact, not a broken kit, and it is knowable before
you run anything. `queries/00-preflight.kql` reports it directly. Run that
first. It tells you, for each table: not present, present but empty,
present without token data, or ready.

You also need read access to the workspace (`Log Analytics Reader` is
sufficient). Nothing here needs write access to anything.

## What you run

`economic-control-check/workbook.json` is an Azure Workbook. Import it,
pick your workspace and a time range, and work down the page.

Every query is also a standalone file under
`economic-control-check/queries/`, and each one runs unchanged when pasted
straight into Log Analytics — they declare their own 30-day window at the
top, which you can edit. The workbook runs the same query bodies with the
time range supplied by its parameter instead. The workbook is generated
from these files, so the two cannot disagree.

| Query | Answers |
|---|---|
| `00-preflight.kql` | Can this workspace answer the question at all? |
| `01-observed-spend.kql` | What is the daily request volume and token consumption? |
| `02-model-distribution.kql` | Where does consumption concentrate, by model and deployment? |
| `03-attribution-gaps.kql` | How much consumption has no accountable workload behind it? |
| `04-retry-and-failure-waste.kql` | How much was consumed by requests that did not succeed? |
| `05-validation-candidates.kql` | Which operations are worth testing before anything changes? |

### How to read the numbers

Two conventions run through every query, and both exist to keep a missing
value from being reported as a real one.

**One row per request, not per log record.** `ApiManagementGatewayLlmLog`
writes request, response and streamed-chunk records separately under a
single `CorrelationId`. Counting records overstates traffic — on a
streaming-heavy estate, by a multiple — so every query collapses the table
to one row per request before it counts or sums anything.

**Disagreement reads as absent.** Where several records report different
values for the same field, the queries return nothing for it rather than
picking one. A `UsageCoveragePct` below 100 means the totals rest on that
share of requests and are a floor, not an estimate of the remainder.

`05-validation-candidates.kql` ranks; it does not conclude. No threshold is
applied, because this kit does not know your quality, latency or
reliability requirements — and a query that invented one would be
recommending a change nobody tested.

## With an AI agent

Hand `AGENT-INSTRUCTIONS.md` to the assistant you already use. It runs in
your environment with your credentials.

## Without an AI agent

Every step is human-followable: `economic-control-check/` holds the
workbook and queries, `metergrade-free/export-apim.md` describes the
export.

## Verifying this kit

The trust chain, in the only direction it works:

    Sigstore identity  authenticates  manifest.json
    manifest.json      authenticates  the verifier and every kit artifact
    the verifier       checks         the complete kit policy

The verifier bundled here is convenience. **Sigstore and the signed
manifest are the authority.** There is no Metergrade key, no Metergrade
endpoint, and nothing to trust that you cannot check with standard tools.

### First, check what this copy actually contains

    ls manifest.json manifest.json.sig manifest.json.pem

A published Metergrade release contains all three. **If the `.sig` and
`.pem` are absent, this copy did not come from the signing release**, and
`./verify` will FAIL it — exit 1, "do not run this kit" — rather than pass
it with a caveat. That is deliberate. An unsigned kit is an unauthenticated
kit, and a verifier that waved one through would be worth nothing.

If you are holding such a copy, its contents can still be checked for
internal consistency with `./verify --no-signature` (exit 3), but nothing
about that check demonstrates Metergrade produced it. Obtain a signed
release instead.

### Normal use

    ./verify              (macOS, Linux)
    verify.cmd            (Windows)

Needs Node 18+ and [cosign](https://github.com/sigstore/cosign). It prints
a report and exits:

| Exit | Meaning |
|---|---|
| `0` | Verified. Signed by the Metergrade release, and the files match the signed manifest. |
| `1` | **Do not run this kit.** It is not what it claims to be. |
| `3` | Contents are internally consistent, but the signature was not checked (`--no-signature`, cosign missing, or an unsigned copy). Nothing here shows Metergrade produced it. |

A missing cosign is a failure, not a skip. A check that turns "could not
run" into success is how a green result comes to mean nothing.

Optional: `./verify --archive path/to/kit.zip` also checks the downloaded
archive against the published `archive_sha256`.

### High assurance

If you are auditing, do not start by running our code. Establish the
manifest first, then use it to establish the verifier, then run it.

**1. cosign authenticates the manifest.** The identity is pinned to the
release workflow. Do not relax it to a wildcard — that accepts a signature
from anyone, and proves only that a signature exists.

    cosign verify-blob \
      --signature manifest.json.sig \
      --certificate manifest.json.pem \
      --certificate-identity-regexp '^https://github\.com/Metergrade/metergrade-platform/\.github/workflows/release-setup-kit\.yml@refs/' \
      --certificate-oidc-issuer 'https://token.actions.githubusercontent.com' \
      manifest.json

**2. The authenticated manifest establishes the verifier.** Compare the
`metergrade-verify.mjs` entry in `manifest.json` with the file on disk:

    sha256sum metergrade-verify.mjs

**3. Now run it.**

    node metergrade-verify.mjs .

Or skip step 3 entirely: `manifest.json` lists every artifact with its
`sha256`, and `checksums.sha256` is the same data in `sha256sum -c` form.
Two things that check does not do for you, and the verifier does — check
that no file is present which the manifest does **not** list (every
declared hash still matches when something undeclared has been added), and
recompute `content_hash` from the files rather than reading it back out of
the manifest.

### What the archive checksum is not

`archive_sha256` in `RELEASE.json` verifies only that the ZIP bytes arrived
intact. It is not the kit's identity — a ZIP embeds timestamps and
ordering, so two honest builds of the same contents produce different ZIP
bytes. `RELEASE.json` is published beside the archive rather than inside
it, since it records that archive's own digest.

`content_hash` in `manifest.json` is the identity. It is a digest over
sorted path and hash pairs, so it reproduces on any machine from the
contents alone.

## Removing it

Delete the imported workbook and, if you enabled diagnostic settings only
for this check, revert them. Nothing else was changed.
