A Gate Is About Deciding What Blocks
One-line summary
The essence of a test gate is not running lots of tests but agreeing in advance on which signals will block a merge and having a machine make that verdict instead.
Why this is needed
The later a defect is caught, the more expensive it gets. The widely cited multiplier from IBM research is 1 in development, 10 in QA and 100 or more in production. Even for the same bug, catching it on my branch takes 30 minutes, but catching it in production adds incident response, a hotfix, a retrospective and even customer communication. A gate is a device placed to catch it at the left end of this cost curve.
The problem is which tests, and how many. The recommended ratio of the test pyramid is 70 unit / 20 integration / 10 E2E. If this gets flipped into an inverted pyramid (an ice cream cone), four things come at once. E2E is slow so the feedback loop gets long, flaky tests are frequent, even when they fail it is hard to tell where the cause is, and the maintenance cost grows exponentially. A state in which there are more and more tests yet nobody trusts the results is exactly this shape.
How it works
A gate is divided into three parts.
First, the criteria for the verdict. Aim for coverage of 80% or more for core business logic and 60–80% overall, but do not forget the principle that tests with meaningful assertions matter more than 100% coverage. Tests that only match the number make only the coverage report pretty and let bugs pass as they are.
Second, the blocking policy. Always block on unit and integration. For E2E, block only on the main flows and leave the rest as observation. Block on performance only when it exceeds a threshold, and on security scans only for CRITICAL. Leave chaos tests non-blocking and only watch the metrics. If you make everything blocking, people first learn how to bypass the gate.
Third, the enforcement point. The place where a gate is actually enforced is not the pipeline script but the required status checks of branch protection. The script only produces the result, and the authority to block a merge belongs to the repository settings. There is the same pattern on the IaC side. Plan is shown in the PR and Apply is run automatically after the merge.
Flaky tests are the main culprit that eats away at the gate's credibility, so they need separate handling. Use explicit waits for a specific condition instead of fixed sleeps, cut the dependencies between tests, mock the network or wrap it in retries with an upper bound, and fix dynamic data such as dates. Retries must always have an upper bound. Infinite retry is not stabilization but hiding failures. For tests that still wobble, put them on a quarantine list so they do not block the merge, but record the fact that they are quarantined in the report so that they are not forgotten.
What it looks like in the field
When you first turn on a gate, a request always comes: "it's urgent right now and this is stopping us from shipping". What is needed then is an exception approval procedure, not a switch to turn the gate off. Another scene you often see is raising the retry count instead of fixing a failing test. If the retry count grows while the failure rate stays the same, that test is already noise, not a signal.
A gate is followed only if it is fast
The biggest reason people bypass a gate is not that the rules are strict but that it is slow. If they must wait 40 minutes until the merge, developers collect small changes and push them up as big chunks, and big chunks are hard to review and hard to revert. Since the purpose of a gate was exactly the opposite, this is not a mere inconvenience but a design failure.
The ways to gain speed are fixed.
- Run only what changed. In a monorepo, look at which packages changed and test only that impact range. Even in a single repository, there is no reason to run the full tests on a commit that changed only documentation.
- Split into stages. Run the fast ones first and report a failure within a few minutes, and put the slow ones after them. If it fails in an earlier stage, the later ones do not even start, so the average wait time drops greatly.
- Run in parallel. Split the tests into several branches and run them at the same time. If tests share state then, parallelization turns straight into instability, so the premise is that each test creates and cleans up its own resources itself.
- Cache. If dependency installation takes several minutes every time, fix that first. It has a bigger effect than cutting test code.
Do not forget to measure time either. If you record how long each pipeline stage takes, you can find the cause when it suddenly becomes slow one day. Without this record, all you hear is "CI has been slow lately", and nobody can answer what became slow and since when.
Finally, there is one cultural condition that decides the success or failure of a gate. It is not leaving a red pipeline unattended. If a day passes with main broken, for every change that comes in afterward it is impossible to tell whether it is that change's fault, and the gate loses its function as a signal completely. You need an agreement that when it is broken, fixing or reverting it takes priority over any other work. Reverting is almost always faster, so even if you do not yet know the cause, the standard approach is to revert first to turn main green and then investigate.
What you will do in the next lab
You build by hand, in shell, a test runner, a coverage gate, retries and quarantine. You build a runner that runs the rest to the end even when something fails and shows the whole picture, a coverage gate that judges exactly at the boundary value, a retry with an upper bound, and a JSON report whose figures and verdict agree.