I renumbered a field and the old client silently read the wrong value
Four kinds of calls and sixteen ways to fail
Summary
gRPC calls come in four kinds — unary, server streaming, client streaming, and bidirectional — and every call ends with one of seventeen status codes. There is no default deadline, so a call can wait indefinitely, and retries work only once you write a policy.
Why this was needed
The previous two modules dealt with the bytes of a single message. But gRPC incidents happen more often at the boundary of a call than inside a message — a call whose response never arrives and nobody cuts off, a call the server says succeeded but the client sees as failed, a retry that sends the same payment twice. The core concepts document states this boundary firmly: "the client and the server each independently determine whether a call succeeded, and their conclusions can differ". The server can conclude "I finished sending the response", and the client can conclude "it arrived after the deadline". Once you accept this sentence, all the remaining rules follow.
How it works
The four kinds of calls. A service definition is written in a .proto with rpc.
service OrderService {
rpc GetOrder (GetOrderRequest) returns (Order); // 단항
rpc ListOrders (ListRequest) returns (stream Order); // 서버 스트리밍
rpc UploadEvents (stream Event) returns (UploadSummary); // 클라이언트 스트리밍
rpc Chat (stream ChatMessage) returns (stream ChatMessage); // 양방향
}
Unary is like a function call — one request, one response. Server streaming receives several responses as a stream for one request, and client streaming is the reverse. In bidirectional streaming the two streams are independent, so both sides decide the order of reading and writing as they like — the server may answer after receiving everything, or it can be ping-pong, receiving one and answering one. What the document guarantees is the order of messages within one call. The order is preserved within each stream, but there is no ordering between the two streams.
Status codes. The status codes document defines seventeen, from 0 to 16. It is important that they divide into those the library generates and those only the application generates. INVALID_ARGUMENT, NOT_FOUND, ALREADY_EXISTS, FAILED_PRECONDITION, ABORTED, OUT_OF_RANGE, and DATA_LOSS are never generated by the library — if you saw one of these codes, it was definitely returned by server code.
| Code | Meaning | Who generates it |
|---|---|---|
| OK 0 | Success | |
| CANCELLED 1 | Canceled by the caller | Library, app |
| INVALID_ARGUMENT 3 | An argument that is wrong regardless of the system state | App only |
| DEADLINE_EXCEEDED 4 | Could not finish within the deadline — it may have finished | Library, app |
| NOT_FOUND 5 | The requested thing does not exist | App only |
| PERMISSION_DENIED 7 | No permission (not used for resource exhaustion) | |
| RESOURCE_EXHAUSTED 8 | Quota or capacity exhausted | |
| FAILED_PRECONDITION 9 | Do not retry until the state is fixed | App only |
| ABORTED 10 | Retry at a higher level (restart the transaction) | App only |
| UNIMPLEMENTED 12 | The method does not exist | |
| INTERNAL 13 | An invariant is broken — reserved for serious errors | |
| UNAVAILABLE 14 | Transient, can be retried after backoff — may not be safe for non-idempotent calls | |
| UNAUTHENTICATED 16 | No authentication credentials |
The guideline for telling three codes apart is right there in the document. If only this call needs to be repeated, UNAVAILABLE; if the read-modify-write sequence must be redone from the beginning, ABORTED; and if it must not be retried until a person fixes the system state, FAILED_PRECONDITION. An rmdir failing on a non-empty directory is the last case. And you must remember one sentence from the explanation of DEADLINE_EXCEEDED — for an operation that changes state, this code can come back even if the operation completed successfully. It may be only that the response arrived late.
The error handling document also gives a table of the situations for codes the library generates. If a server handler throws an exception, UNKNOWN; if the method does not exist, UNIMPLEMENTED; if the server is going down, UNAVAILABLE; and if it cannot parse the request protobuf, INTERNAL. The standard error model has only a code and a string message, and if you need structured detail, you use the extended model that carries google.rpc.Status in the trailers — though the cost is noted that proxies and loggers cannot see inside it and HTTP/2 header compression efficiency drops.
Deadlines. The first rule of the deadlines document is "gRPC does not set a default deadline". If the client does not set one, it can wait forever, so it says always to set a realistic value. A deadline is a point in time and a timeout is a duration; the API differs by language but the meaning is the same. When the deadline passes, the client fails the call with DEADLINE_EXCEEDED, and the server automatically cancels that call (CANCELLED). But it is the server code's responsibility for the server application to stop what it was doing — the library has no way to interrupt a handler, so long operations must periodically check whether they have been canceled.
Propagation is the key. When your server calls another server, it must pass on the original client's deadline. The document notes that Java and Go propagate by default and C++ must be turned on. If the point in time were passed on as is, the clocks of the two servers could be out of step, so gRPC converts it into a timeout with the elapsed time already subtracted and passes that. The document's example is exactly this picture — if the client gave 2 seconds and the user server spent 0.5 seconds and then called the billing server, the billing server receives 1.5 seconds.
Cancellation. According to the cancellation document, a client can signal at any time that it has lost interest, and deadline expiry and I/O errors also cause cancellation. Ideally cancellation propagates upstream, so Java, Go, and C++ automatically cancel outgoing calls. One warning line from the core concepts document — what has already changed before the cancellation is not rolled back.
Retries. The retry document clarifies the part with the most misunderstanding. Retries are on by default, but there is no default policy. Without a policy, gRPC does only "transparent retries" — unlimited if the call never left the client, and exactly once if it reached the server library but the application logic did not see it. If there is a chance the server processed it, it does not retry. The moment response headers are received, the call is committed and is no longer retried.
The policy is written per method in the service config.
"retryPolicy": {
"maxAttempts": 4,
"initialBackoff": "0.1s",
"maxBackoff": "1s",
"backoffMultiplier": 2,
"retryableStatusCodes": ["UNAVAILABLE"]
}
Backoff has ±20% jitter attached, so an initial 0.1 seconds actually becomes between 80–120ms. So that retries do not knock the server down again, there is retryThrottling (maxTokens and tokenRatio), where each failure reduces the tokens by 1, each success increases them by tokenRatio, and retries stop when they fall below half. And the single line in the UNAVAILABLE explanation of the status codes document is the whole of policy design — retries may not be safe for non-idempotent operations. To put a call such as payment creation in retryableStatusCodes, the server must be preventing duplicates with an idempotency key.
Health checking and metadata. The health checking document defines the standard service health/v1. The unary Check is for central monitoring, and the streaming Watch is for a client to attach and receive state changes. It reports SERVING or NOT_SERVING per service name, and the empty string means the whole server. If the client turns on healthCheckConfig, it does not send requests until Watch says it is healthy, and stops sending when it becomes unhealthy. If Watch fails with UNIMPLEMENTED, health checking is turned off. The metadata document describes key-value pairs carried as HTTP/2 headers — keys are ASCII, case-insensitive, cannot start with grpc-, and keys for binary values end with -bin. Headers go before the first message, and trailers are sent when the server closes the call.
What you meet in the field
The most expensive incident is a call without a deadline. When a downstream service stopped, all the upstream threads piled up waiting for responses, and eventually even the upstream died. The cause was not knowing the sentence that gRPC does not set a default deadline. If you give a deadline and propagate it, even when the downstream is late, the upstream returns within the set time with DEADLINE_EXCEEDED.
The second is "succeeded but failed". An order creation was committed on the server, but the response exceeded the deadline so the client received DEADLINE_EXCEEDED, and because the retry policy contained that code, the same order was created twice. This is exactly the case the document records in the DEADLINE_EXCEEDED explanation as "it may have finished". You must not widen the retry list without an idempotency key.
The third is a health check misunderstanding. The load balancer called Check but wrote the service name wrongly and always received NOT_SERVING, and traffic was 0. It is a setting that ends in one line once you know that the empty string means the whole server.
What to check in the next quiz
It asks about the differences among the four kinds of calls, the status codes the library never generates, telling UNAVAILABLE, ABORTED, and FAILED_PRECONDITION apart, how a deadline avoids clock skew, what actually happens when there is no retry policy, and the empty string in health checks and the metadata key rules.