// @title Observed consumption by day // @summary Request volume and token consumption per day, collapsed to one row per request. Usage coverage states how much of the volume the totals rest on. // @posture count // @posture-note Aggregates observed consumption. No comparison, no classification. // // Observed AI request volume and token consumption, by day. // Read-only. Runs in your Log Analytics workspace; sends nothing anywhere. // // ONE ROW PER REQUEST, NOT PER LOG RECORD. // // ApiManagementGatewayLlmLog does not emit one record per request. Request, // response and streamed-chunk records are written separately under a single // CorrelationId. Counting records therefore overstates traffic — on a // streaming-heavy estate, by a multiple — and summing token columns across // those records double-counts any value reported more than once. // // So the table is collapsed to one row per CorrelationId first, and a token // value is taken only when the records AGREE on it: make_set(x, 2) collects // up to two distinct non-null values, so array_length() == 1 proves // agreement, == 0 means nothing was reported, and > 1 is a real conflict. // A conflict yields null — Metergrade declines the number rather than // choosing one. Summing null is not summing zero, which is why the coverage // column below exists: it states how much of the volume the totals rest on. // // A wrong number is worse than a missing one. let _startTime = ago(30d); let _endTime = now(); let requests = ApiManagementGatewayLlmLog | where TimeGenerated between (_startTime .. _endTime) | summarize TimeGenerated = min(TimeGenerated), PromptTokensSet = make_set(PromptTokens, 2), CompletionTokensSet = make_set(CompletionTokens, 2), TotalTokensSet = make_set(TotalTokens, 2) by CorrelationId | extend PromptTokens = iff(array_length(PromptTokensSet) == 1, tolong(PromptTokensSet[0]), long(null)), CompletionTokens = iff(array_length(CompletionTokensSet) == 1, tolong(CompletionTokensSet[0]), long(null)), TotalTokens = iff(array_length(TotalTokensSet) == 1, tolong(TotalTokensSet[0]), long(null)); requests | summarize Requests = count(), RequestsWithUsage = countif(isnotnull(TotalTokens)), InputTokens = sum(PromptTokens), OutputTokens = sum(CompletionTokens), TotalTokens = sum(TotalTokens) by Day = bin(TimeGenerated, 1d) // The share of requests the token totals are actually derived from. Below // 100%, the totals are a floor and not an estimate of the remainder. | extend UsageCoveragePct = round(100.0 * RequestsWithUsage / Requests, 1) // // KUSTO'S sum() RETURNS 0 FOR A GROUP WHERE EVERY VALUE IS NULL, not null. // That turns "no request in this group reported usage" into the confident // claim "this group consumed nothing" — the exact substitution the rest of // the kit refuses to make. A bar at zero reads as measured absence of spend; // a gap reads as absence of evidence, which is what it is. So a token total // resting on zero reporting requests is returned absent. | extend InputTokens = iff(RequestsWithUsage > 0, InputTokens, long(null)), OutputTokens = iff(RequestsWithUsage > 0, OutputTokens, long(null)), TotalTokens = iff(RequestsWithUsage > 0, TotalTokens, long(null)) | project Day, Requests, RequestsWithUsage, UsageCoveragePct, InputTokens, OutputTokens, TotalTokens | order by Day asc