errno

Two Azure CLI traps: -o tsv smuggles a carriage return, and az monitor metrics list returns Bad Request

· in Servers and infrastructure · tested on Azure CLI on Linux, Azure Monitor REST API 2018-01-01

Symptom

I was writing a small reporting script: pull a couple of values out of Azure with az, then use them to build further calls. Nothing exotic — a resource id here, an API version there, a metric query at the end. Two separate things went wrong, and both cost far more time than they should have, because both hid the real problem behind an error message about something else.

The first one looked like a service-side complaint. A value captured from az ... -o tsv went into a variable, the variable went into a query string as ?api-version=$V, and the service answered with NoRegisteredProviderFound — and quoted back an api-version that was not the one I had asked for. The string in the message looked like my value, plus something at the end that the terminal would not show me.

The second one was blunt and detail-free:

Operation returned an invalid status 'Bad Request'

That is the whole error. It came from az monitor metrics list, for metrics that exist and that I can see in the portal: UsedCapacity on a storage account, Network Out Total on a virtual machine. No indication of which parameter the service objected to, no correlation id in the output, nothing to grep for.

The wrong diagnosis, and how to rule it out

For the first trap the obvious reading of NoRegisteredProviderFound is that the resource provider really is not registered in the subscription, or that the api-version I picked is not supported for that provider. Both are real conditions and both have well-known fixes, so it is easy to spend twenty minutes registering providers and reading the supported-versions list.

The tell that it is neither: the api-version in the error message is not the api-version you typed. Read that string character by character. If the service echoes back something that is your value plus a mangled tail, the value never arrived intact — and then nothing about providers is relevant.

For the second trap the instinct is to blame the metric name, the aggregation, the time range, or permissions. Permissions are easy to eliminate: the same identity can list the resource and read other metrics in the portal. The metric name is easy to eliminate too — the metric definitions endpoint will happily tell you the metric exists. What remains is a parameter the CLI is marshalling wrongly, and since the error carries no parameter name, guessing your way through the flag combinations is a search with no feedback.

What I measured

The first trap is a carriage return. az ... -o tsv can hand you a value with a trailing \r. A \r returns the cursor to the start of the line, so echo "$V" prints something that looks exactly right, and in a longer line it will happily overwrite what came before it. It is invisible in every place you would normally look.

Make it visible at the byte level:

V=$(az provider show --namespace Microsoft.Insights --query "resourceTypes[0].apiVersions[0]" -o tsv)

printf '%s' "$V" | od -c        # trailing \r shows as \r
printf '%s' "$V" | cat -A       # trailing \r shows as ^M, line end as $
printf '%q\n' "$V"              # shell-quoted form, escapes the CR

That is a two-second check and it ends the argument. The comparison that fails for no visible reason — [ "$V" = "2018-01-01" ] returning false while both sides print identically — is the same fact seen from a different angle.

For the second trap what I measured was simply that the same query works one layer down. The metric, the resource, the timespan, the aggregation: all unchanged, issued as a plain REST call, and the data comes back. That isolates the bug to the CLI’s command layer rather than to the request I was trying to make.

The fix

Strip the carriage return at the point of capture. Every -o tsv capture, no exceptions:

V=$(az provider show --namespace Microsoft.Insights \
      --query "resourceTypes[0].apiVersions[0]" -o tsv | tr -d '\r')

If you have more than a handful of these, put the normalisation in one place so it cannot be forgotten:

azv() { az "$@" -o tsv | tr -d '\r'; }

V=$(azv provider show --namespace Microsoft.Insights --query "resourceTypes[0].apiVersions[0]")

Better still, stop using tsv for anything structured. -o json piped through jq -r does not have this problem at all, and it survives values containing tabs, which tsv does not:

V=$(az provider show --namespace Microsoft.Insights -o json \
      | jq -r '.resourceTypes[0].apiVersions[0]')

tsv is fine for a human reading one field on a terminal. For a value you are about to interpolate into a URL, a filename or a comparison, treat it as untrusted text.

For the metrics call, drop to the REST endpoint through the CLI itself. az rest sends an arbitrary request with the credential you are already logged in with:

az rest --method get --url "https://management.azure.com/subscriptions/<sub>/resourceGroups/<rg>/providers/Microsoft.Storage/storageAccounts/<acct>/providers/Microsoft.Insights/metrics?api-version=2018-01-01&metricnames=UsedCapacity&timespan=<start>/<end>&interval=P1D&aggregation=total"

Same metric, same account, same identity — and it returns the time series. You keep the part of the CLI that is genuinely hard to replace (token acquisition, refresh, tenant and cloud selection) and skip the part that was broken (parameter marshalling in one command).

Two things to get right in that URL. timespan is a pair of ISO 8601 timestamps separated by a slash; anything else is rejected, and this is one of the few cases where the service does tell you what it disliked. And aggregation must be one the metric actually supports — total, average, minimum, maximum, count are not interchangeable, and asking for an unsupported one on a given metric gets you nothing useful back. If you are unsure, the metric definitions endpoint for the same resource lists the supported aggregations per metric, and that list is the authority, not the portal’s chart menu.

The transferable lesson

Both of these bugs live in the layer between me and the API. One is the shell capture: the bytes I thought I had were not the bytes I had. One is the CLI’s own translation of flags into an HTTP request: the command I thought I was sending was not the request that went out. Neither error message pointed at the layer that was actually wrong, and in both cases the service was behaving correctly given what it received.

So the habit is to make that boundary visible rather than reason about it. When a value misbehaves, dump it at the byte level — od -c, cat -A, printf '%q' — before you theorise about the remote end. When a wrapper misbehaves, go one layer down to the documented HTTP call and see whether the request itself is fine; if it is, you have found the bug and you also have a working workaround in the same step.

There is a practical reason to reach for those checks early with Azure specifically: az commands are slow. A few seconds of startup and a round trip per invocation means a guess-and-retry loop costs minutes, not milliseconds. A byte-level look at what you are actually sending is cheaper than one more guess, and unlike the guess, it either ends the investigation or definitively eliminates a whole class of cause.

azure cli bash monitoring