TT Lab
Get started
Learn Learning paths Courses

CAPA — Argo Project Associate

One order, two workflow executions: why?

Continue in TT Lab

In one line

That an event was received, that a workflow ran, and that an order was reflected in the ledger are different pieces of evidence. If you treat all three as the same success, you can miss failures and create duplicate processing.

Why this was needed

Imagine you run a fictional order intake server. A user pressed the order button once but could not see the response because the network was unstable. The screen tells them to try again, and the user sends the same order once more. Two HTTP requests arrive at the server. In an event system that handles each of the two requests as a new event, separate runs can be created. The user experience of pressing the button once does not guarantee a single run on the server.

There is also a misunderstanding in the opposite direction. The webhook responded with 200, but the business did not start. A filter may have rejected the input, or the Sensor may lack permission to create the Workflow. In this case, if you report success looking only at the client response, the next person investigating the failure starts out on a wrong assumption from the beginning.

This lab is not a real payment but a small order ledger. You deliberately apply a wrong path, send unexpected input, and leave out the submit permission. Afterwards you send the same order twice and compare the number of times the business effect was applied. The goal is not to experience failure but to learn how to narrow down the failed layer with evidence.

How it works

The ready signal has boundaries too

The EventSource receives HTTP requests and sends them to the event bus, and the Sensor subscribes to them and evaluates conditions. A Deployment being ready is a Pod readiness condition; it does not mean that the subscription to a particular event is all finished. In the probe, even the normal control group sent right after the Pod became ready did not run. This version's JetStream connection implementation targets messages after the new consumer starts subscribing. So before sending the first event, you must check the relevant Sensor Pod and the subscription state together.

You must also read the tense of the logs. Subscribing means an attempt is in progress and is recorded before the subscribe call. After checking the log of a successful subscription and the actual consumer waiting state, you send the first business event. Simply sleeping a few more seconds can fail again in a slow environment. This check is a procedure to align the starting point of the lab and is not a guarantee of preventing every event loss in a production environment.

A filter is part of the input contract

The HTTP body of a webhook event goes into body inside the event data. If the body has kind, the data filter path is body.kind. If you wrote body.event.kind, the HTTP body must have a nested object called event. To see the rejection when the path is missing, the normal control group that the same Sensor should accept must also succeed. From the fact that no workflow exists alone, you cannot conclude the filter is accurate.

A string filter is evaluated as a regular expression. The dot in the pattern order.created is not a literal dot, and the start and end are not anchored. In the measurement, pre-orderXcreated-tail also passed. To allow exactly one kind, you must treat the dot as a character and constrain the whole string. A change that widens the allowed set as inputs increase should be checked with special care.

type: bool is also not the same as strict JSON schema validation. In this lab's version, not only the boolean true but also the string true and the number 1 passed. If only a JSON boolean should be allowed, you must check the type of the original value, not the converted result. In the lab you tell them apart with a Lua filter. This is a type contract for a specific field and does not replace validating the whole request schema.

When you combine several conditions, OR passes if any one is satisfied and AND passes only if all are satisfied. For a business that must not run just because enabled is true even when the kind is wrong, you should review the operator between the two. Also do not assume that the error of a nonexistent field, which is just the rejection of one condition, invalidates the approval of another condition joined by OR.

Who creates and who executes

Separate the identity by which the Sensor creates the Workflow from the identity used by the tasks inside the Workflow. This Kubernetes resource creation trigger allows the submit account only Workflow create. The business account has the permission to record the needed task results. If you attach an administrator role to both accounts to get rid of the permission error, you do not learn which permissions are actually needed and the blast radius also widens.

The HTTP response for a rejected event can be 200. So you look together at the Sensor's actual forbidden message, the request's trace ID, and whether a Workflow was created. After fixing with minimal permissions, you confirm the recovery with a new event. To claim that a past event that already failed is automatically reprocessed, you must separately observe that past ID.

Arrival ID, execution ID, and business key

It is not strange for two HTTP requests for the same order to get different arrival IDs. The fact that the two Workflow UIDs are different also does not mean two orders. The business key must point to the same order even on retry. This ledger uses the order ID as the business key, and ties together checking whether the key was already applied and applying the amount in a single database transaction.

In the naive implementation, both runs add the amount. In the improved implementation, both runs can end normally, but only one changes the ledger and the other treats it as a duplicate. If the key is the same but the amount differs, it is not hidden as a normal retry and is rejected as a conflict. Even if the success response could not be read, if the committed record remains, the next request may not increase the business effect.

What it looks like in the field

If you put only one success count on a business dashboard, it is hard to tell which success it is. Think of it as split into the number of HTTP intakes, the number of approved events, the number of completed Workflows, and the number of business applications. During a time when there were duplicate requests, these numbers can differ while the system still behaves as intended. Conversely, even if they are all the same number, if wrong inputs were all approved, the business is wrong.

This single-VM ledger is not a high-availability database or a real payment system. If you send an email or call a payment outside the same SQLite transaction, those external effects are not atomically tied together. For such requirements, you must consider a separate delivery contract such as an external service's idempotency key or an outbox. What was confirmed here is this lab's local ledger effect, and it does not guarantee retention after the session ends or a single execution of every external side effect.

What you will do in the next lab

You break the filter path, regular expression, type, and logical operation and compare with the normal control group. After fixing a real permission denial with a minimal Role, you check, in a duplicate request, the two executions and the one business effect separately. At the end, you leave a short incident report that distinguishes which layer you observed and what has not yet been proven.

Official references: Data filter, Script filter, Service accounts, Pinned-version subscription implementation, SQLite transactions.