We raised the shard count and the failing test started passing
Goal
You split a test suite deterministically, leave per-piece results in a machine-readable format, and merge them again into a single verdict. You also build the check that protects against the number of pieces changing the verdict, and an allocation that reduces the heaviest piece using recorded durations.
Why it matters
When tests get long, people do not wait for the result. Running them split is the most direct means of reducing wall-clock time without deleting tests, but splitting brings in a new failure mode with it. If the allocation differs from run to run, a record saying 'failed in piece 2' points to a different test in the next run, making reproduction impossible. If you merge per-piece results carelessly, the whole looks green even though an entire piece died. And the order dependencies that hid by always running in the same order are exposed the moment the pieces diverge. So a pipeline that has introduced splitting needs one invariant — even if you change the number of pieces, the total case count and the overall verdict must be the same. This lab has you create the scene where that invariant breaks, find the cause and fix it, and then build the gate that keeps it from breaking again.
Steps
- Create the library under test,
/root/shard/tests/sampleapp.py(it must have a module-level settings dictionarySETTINGS), and in the same directory four or moretest_*.pyfiles and twenty or more cases. One of them is/root/shard/tests/test_slowio.py, which has five cases that take 0.1 seconds or more withtime.sleep(this will create the time difference between pieces later). Every test in this step must pass. And create/root/shard/list_tests.sh [시험디렉터리](optional test directory) so that it outputs the test names in the form모듈.클래스.메서드(module.class.method), sorted, one per line (the default is/root/shard/tests). Finally, save the output of running everything at once to/root/shard/baseline.txt(it must containRan N testsandOK). - Create
/root/shard/shard.sh <조각번호> <조각수>(piece number, number of pieces). From the list of test names received on standard input, it outputs as they are only the names that belong to that piece. The piece is decided by the remainder of the first 8 hexadecimal digits of the name's sha256 divided by the number of pieces. Piece numbers count from 0. Three things must hold — (1) the pieces do not overlap and all together make up the whole input, (2) the same input gives output that is identical down to a single character even when run twice, (3) removing one test from the list does not change the piece of the remaining names. Leave the confirmed result in/root/shard/deterministic.txt— one line saying that thediffof the two runs' output is empty, and one line with the counts when split into 3 pieces in the formshard0=<수> shard1=<수> shard2=<수>(count). - Create
/root/shard/runner.py <시험디렉터리> <시험목록파일> <출력XML> <조각이름>(test directory, test list file, output XML, piece name). It runs the tests in the list in a single process in list order and writes JUnit XML. The root element istestsuite(with the attributesname,tests,failures,errorsandtime), each test gets atestcase(with the attributesclassname,nameandtime), and a failure is left as afailurechild element and an error as anerrorchild element. And create/root/shard/run_shard.sh <조각번호> <조각수> [리포트디렉터리](piece number, number of pieces, optional report directory) so that it makes the list, splits it and runs it, leaving<리포트디렉터리>/shard-<조각번호>.xml(the report directory default is/root/shard/reports). Create the reports of the three pieces withbash /root/shard/run_shard.sh 0 3,1 3and2 3. - Create
/root/shard/merge.sh <리포트디렉터리> <기대조각수>(report directory, expected number of pieces). It reads all the*.xmlin that directory and prints just one line:PASS cases=<수> failures=<수> errors=<수> shards=<수>(counts) or aFAIL ...of the same shape. If even one of four things is off, it isFAILand the exit code is 1 — the number of report files differs from the expected number of pieces, there is an XML that cannot be read, there is not a single case, or there is even one failure or error. Get the counts not from the header's attribute values but by counting the actualtestcase,failureanderrorelements. On the/root/shard/reportsmade in step 3, runmerge.sh /root/shard/reports 3and save its output to/root/shard/verdict.txt. - Create
/root/shard/total_of.sh <조각수> <출력디렉터리>(number of pieces, output directory). It runs all the pieces with that number of pieces, leaves the reports in the output directory, and at the end prints the one line ofmerge.shas it is (passing along the exit code too). The reports left in the output directory must be exactly as many as the number of pieces, and reports from earlier runs must not get mixed in. Runtotal_of.sh 2 /root/shard/out2andtotal_of.sh 3 /root/shard/out3and save the two lines to/root/shard/invariant.txt. What remains after removingshards=(cases=,failures=anderrors=) must be equal between them. - You deliberately put in tests that lean on shared state.
/root/shard/tests/test_config.pyhas two cases that change the value ofsampleapp.SETTINGSfor testing but do not restore it, and/root/shard/tests/test_profile.pyhas two cases that expect the default settings as they are. Confirm that if you run everything in one process, the later tests fail, and that if you run only those tests alone, they pass. Then runtotal_of.shwith 2, 3, 4 and 5 pieces to see the overall verdict split depending on the number of pieces, and write the id of the test that did the contaminating as the first line and the id of the test that suffered as the second line in/root/shard/order-dep.txt. Finally, fix both tests so that they keep usingsampleapp.SETTINGSand yet pass in whatever order they are run (do not avoid it by deleting the tests or by not using the shared settings). - Create
/root/shard/timings.sh <리포트디렉터리>(report directory) so that, from thetestcaseelements of the reports, it outputs<시험id> <초>(test id, seconds) in name order. With it, build the record for all the tests now at/root/shard/timings.txt(you can use a report from running everything as one piece). Next create/root/shard/balance.sh <시간기록파일> <조각수>(timing record file, number of pieces) — it splits using the method of putting the longest first into whichever piece is lightest at that moment (LPT) and prints<조각번호> <시험id>(piece number, test id) one per line (every test must appear exactly once). Finally, for 3 pieces, measure the sum of the heaviest piece for the hash allocation of step 2 and for this balanced allocation, and write two lines in/root/shard/balance.txt:equal <초>andbalanced <초>(seconds). - Create
/root/shard/order_fast_first.sh <시간기록파일> <시험목록파일>(timing record file, test list file). It outputs the list in ascending order of recorded duration (by name if equal), and a test with no record is treated as 0 seconds and put at the very front. Next create/root/shard/gate.sh <조각수> [시험디렉터리](number of pieces, optional test directory) (the test directory default is/root/shard/tests). It makes the list, splits it into pieces, runs each piece fastest first, and merges, outputting the one line ofmerge.shand its exit code as they are. It leaves reports only in its own temporary directory and leaves no XML in/root/shard. Save the output ofbash /root/shard/gate.sh 3to/root/shard/gate.txt.
Notes
- The work directory is unified as
/root/shard. Every tool the student builds is designed to take the report directory or the test directory as an argument — that way you can run it again on a temporary copy. - Containers cannot be started in this Pod (seccomp blocks user namespaces). Do not use
podman run,podman buildorbuildah. What you will use is bash, the python3.12 standard library, jq and sha256sum.bc,makeandpytestare not available. - Common mistake: splitting by line number or array index of the list. If one test is added or removed, everything after it shifts and the record 'failed in piece N' becomes meaningless.
- Common mistake: the merge not counting the number of report files. If one piece dies, there is no file at all, so looking only at the remaining reports, everything is green.
- Common mistake: trusting only the
failuresattribute in the JUnit XML header. That value is a number written by whoever wrote the report, so it is safer to count the elements directly. - Continuous Integration (10-minute build) · GitHub Actions: splitting jobs with a matrix · GitLab CI: job artifacts
Build the test suite and run it all at once to set a baseline
Create the library under test, /root/shard/tests/sampleapp.py (it must have a module-level settings dictionary SETTINGS), and in the same directory four or more test_*.py files and twenty or more cases. One of them is /root/shard/tests/test_slowio.py, which has five cases that take 0.1 seconds or more with time.sleep (this will create the time difference between pieces later). Every test in this step must pass. And create /root/shard/list_tests.sh [시험디렉터리] (optional test directory) so that it outputs the test names in the form 모듈.클래스.메서드 (module.class.method), sorted, one per line (the default is /root/shard/tests). Finally, save the output of running everything at once to /root/shard/baseline.txt (it must contain Ran N tests and OK).
unittest's defaultTestLoader.discover(디렉터리, top_level_dir=디렉터리) gives you the test suite, and the id() of each test object is exactly the 모듈.클래스.메서드 string. The list is the input of all the later steps, so output it sorted and without duplicates. The full run is done with cd tests && python3 -m unittest $(목록).
Make the same name always go to the same piece
Create /root/shard/shard.sh <조각번호> <조각수> (piece number, number of pieces). From the list of test names received on standard input, it outputs as they are only the names that belong to that piece. The piece is decided by the remainder of the first 8 hexadecimal digits of the name's sha256 divided by the number of pieces. Piece numbers count from 0. Three things must hold — (1) the pieces do not overlap and all together make up the whole input, (2) the same input gives output that is identical down to a single character even when run twice, (3) removing one test from the list does not change the piece of the remaining names. Leave the confirmed result in /root/shard/deterministic.txt — one line saying that the diff of the two runs' output is empty, and one line with the counts when split into 3 pieces in the form shard0=<수> shard1=<수> shard2=<수> (count).
Get eight hexadecimal digits with printf '%s' "$name" | sha256sum | cut -c1-8, and get the remainder with $(( 16#$h % n )) in bash arithmetic. If you split by line number or array index, (3) breaks — if even one is removed at the front, everything after it shifts. Use read as IFS= read -r to preserve spaces.
Run a piece and leave a machine-readable result
Create /root/shard/runner.py <시험디렉터리> <시험목록파일> <출력XML> <조각이름> (test directory, test list file, output XML, piece name). It runs the tests in the list in a single process in list order and writes JUnit XML. The root element is testsuite (with the attributes name, tests, failures, errors and time), each test gets a testcase (with the attributes classname, name and time), and a failure is left as a failure child element and an error as an error child element. And create /root/shard/run_shard.sh <조각번호> <조각수> [리포트디렉터리] (piece number, number of pieces, optional report directory) so that it makes the list, splits it and runs it, leaving <리포트디렉터리>/shard-<조각번호>.xml (the report directory default is /root/shard/reports). Create the reports of the three pieces with bash /root/shard/run_shard.sh 0 3, 1 3 and 2 3.
With a unittest.TestResult(), a suite built with loadTestsFromName(id) can be run(), and you can receive the result of a single case separately, and if you measure time.time() before and after it, you get the per-case duration. Build the XML with xml.etree.ElementTree. You must put the test directory into sys.path for the 모듈.클래스.메서드 names to resolve. Running them in succession in one process is important — the difference shows in step 6.
There is no verdict before merging
Create /root/shard/merge.sh <리포트디렉터리> <기대조각수> (report directory, expected number of pieces). It reads all the *.xml in that directory and prints just one line: PASS cases=<수> failures=<수> errors=<수> shards=<수> (counts) or a FAIL ... of the same shape. If even one of four things is off, it is FAIL and the exit code is 1 — the number of report files differs from the expected number of pieces, there is an XML that cannot be read, there is not a single case, or there is even one failure or error. Get the counts not from the header's attribute values but by counting the actual testcase, failure and error elements. On the /root/shard/reports made in step 3, run merge.sh /root/shard/reports 3 and save its output to /root/shard/verdict.txt.
If an entire piece dies, the report file itself is missing. That is why you must receive 'how many were expected' as an argument. The failures attribute in the header is a value written by whoever wrote the report, so if you trust only that, a report with only the header rewritten to 0 passes. Walk the elements directly with ET.parse(f).getroot().iter("testcase").
Even if you change the number of pieces, the overall verdict must be the same
Create /root/shard/total_of.sh <조각수> <출력디렉터리> (number of pieces, output directory). It runs all the pieces with that number of pieces, leaves the reports in the output directory, and at the end prints the one line of merge.sh as it is (passing along the exit code too). The reports left in the output directory must be exactly as many as the number of pieces, and reports from earlier runs must not get mixed in. Run total_of.sh 2 /root/shard/out2 and total_of.sh 3 /root/shard/out3 and save the two lines to /root/shard/invariant.txt. What remains after removing shards= (cases=, failures= and errors=) must be equal between them.
The number of pieces is only a way of dividing the work, so it must not affect the total case count or the verdict. If the two lines differ, the allocator either dropped tests or is counting some twice. If XML from an earlier run remains, the counts swell, so empty the output directory at the start.
A test that passed alone collapses when run together
You deliberately put in tests that lean on shared state. /root/shard/tests/test_config.py has two cases that change the value of sampleapp.SETTINGS for testing but do not restore it, and /root/shard/tests/test_profile.py has two cases that expect the default settings as they are. Confirm that if you run everything in one process, the later tests fail, and that if you run only those tests alone, they pass. Then run total_of.sh with 2, 3, 4 and 5 pieces to see the overall verdict split depending on the number of pieces, and write the id of the test that did the contaminating as the first line and the id of the test that suffered as the second line in /root/shard/order-dep.txt. Finally, fix both tests so that they keep using sampleapp.SETTINGS and yet pass in whatever order they are run (do not avoid it by deleting the tests or by not using the shared settings).
If you run in succession in one process, module-level variables stay as the earlier test changed them. The place to fix is one of two — the side that changes it restores it in tearDown, or the side that expects sets up the value it will use in setUp. test_render.py already shows the right way. Also check that the alphabetical order of file names is the execution order.
Splitting again by recorded times made the heaviest piece lighter
Create /root/shard/timings.sh <리포트디렉터리> (report directory) so that, from the testcase elements of the reports, it outputs <시험id> <초> (test id, seconds) in name order. With it, build the record for all the tests now at /root/shard/timings.txt (you can use a report from running everything as one piece). Next create /root/shard/balance.sh <시간기록파일> <조각수> (timing record file, number of pieces) — it splits using the method of putting the longest first into whichever piece is lightest at that moment (LPT) and prints <조각번호> <시험id> (piece number, test id) one per line (every test must appear exactly once). Finally, for 3 pieces, measure the sum of the heaviest piece for the hash allocation of step 2 and for this balanced allocation, and write two lines in /root/shard/balance.txt: equal <초> and balanced <초> (seconds).
The timing record is an artifact that becomes the input of the next run — it is already in the reports, so extract and use it rather than measuring anew. LPT sorts the list in descending duration and puts each into the piece whose current sum is smallest. This Pod has no bc, so get the sums with awk or python3.
Bundle splitting, running and merging into one command
Create /root/shard/order_fast_first.sh <시간기록파일> <시험목록파일> (timing record file, test list file). It outputs the list in ascending order of recorded duration (by name if equal), and a test with no record is treated as 0 seconds and put at the very front. Next create /root/shard/gate.sh <조각수> [시험디렉터리] (number of pieces, optional test directory) (the test directory default is /root/shard/tests). It makes the list, splits it into pieces, runs each piece fastest first, and merges, outputting the one line of merge.sh and its exit code as they are. It leaves reports only in its own temporary directory and leaves no XML in /root/shard. Save the output of bash /root/shard/gate.sh 3 to /root/shard/gate.txt.
This is the problem of chaining together what you built in the earlier steps — list_tests.sh, shard.sh, order_fast_first.sh, runner.py and merge.sh. Create the temporary directory with mktemp -d and delete it with trap ... EXIT. The reason it takes the test directory as an argument is to check that the gate really shows red on a copy in which you have deliberately broken a test.