The Sensor Is Fine. The Board Cannot Understand It
I Do Not Even Know Which Bus This File Is
Goal
Decode four I2C captures and three SPI captures yourself, and at the end receive three captures with no names attached, sort out which bus each is, and report the faults with numeric evidence attached.
Why it matters
The one line "the device is not responding" that a driver puts out lumps together four different causes. The address is wrong, the slave is buying time, the pull-up is weak so the rise uses up the whole bit time, or the bus is locked up entirely. These four look completely different on a waveform, and the people who fix them differ too — firmware, the driver's recovery procedure, and hardware. If you sort them out from the waveform, one line of the report points to one person in charge. SPI, conversely, comes with its clock, so there is no speed problem, and instead which edge you read on and when the select line falls create quiet faults.
Steps
- Read a two-line capture — Measure the sample rate, the rise time, and the distribution of the SCL low intervals.
- Find START and STOP — Count only SDA changes while SCL is high as conditions.
- Read out the address and ACK — Recover the bytes and the ACKs from groups of 9 bits.
- An address that no one responded to — Confirm what a NACK is actually the absence of.
- An interval the master did not make — Find clock stretching from the length distribution.
- The rise uses up the bit time — Prove a weak pull-up with the rise time and two thresholds.
- Which of the four combinations — Determine the mode from the idle level and the edge where the data changes.
- If the select line is late, the count tells you first — See the multiple of 8 break.
- Sort out three captures with no names — Report the bus and the fault together with evidence.
Notes
- The materials are under
/opt/fixtures/serialbus/. They arei2c/read-ok.csv,i2c/nack.csv,i2c/stretch.csv,i2c/weak-pullup.csv,spi/mode0.csv,spi/mode3.csv,spi/cs-late.csv, andunknown/capture-1.csvthroughcapture-3.csv. The format description is/opt/fixtures/serialbus/README.md. - The channel names differ from file to file. For I2C they are
sclandsda, for SPIcs,sclk,mosi, andmiso, and for captures of unknown identity they start fromch0. Do not expect meaning in the names; read them from the header. - Common mistake one: when looking for SDA changes while SCL is high, counting even samples that overlap with an edge. It is a condition only when both
scl[i]andscl[i-1]are high. - Common mistake two: reporting the NACK of the last byte as a fault. In a read transfer it is the normal termination procedure.
- The supply voltage is 3.30 V, and the default decision threshold is half of that, 1.65 V. Only in the step that tells you to use VIH do you use 0.7 times.
- No extra installation or internet is needed. You use only python3.
- The expected time is 85 minutes, so extend with +time before the default 60 minutes end (maximum 180 minutes). When the session ends,
/rootdisappears, so keep separately what you want to keep.
Read a two-line capture
Create a working folder with mkdir -p /root/bus-triage, read /opt/fixtures/serialbus/i2c/read-ok.csv, and save to /root/bus-triage/levels.json sample_rate_hz, sample_count, duration_us, channels (the list of channel names in the header), scl_rise_time_samples, sda_rise_time_samples, scl_low_runs, and scl_low_median_us. The rise time is the number of samples of the last complete rise in the capture file (from 0.33 V or below to 2.97 V or above), and an SCL low interval runs from the sample where it went down to the sample where it came back up.
0.33 V and 2.97 V are 10 % and 90 % of 3.30 V. A low interval is between a falling edge and the next rising edge, and an interval that ends still low at the very end has an unknown length, so do not count it.
Find START and STOP
Implement conditions(scl_v, sda_v, threshold_v=1.65) in /root/bus-triage/i2c.py. While SCL is high for two consecutive samples, if SDA falls, put ("start", i) in the list, and if it rises, ("stop", i), and return them in time order. An SDA change while SCL is low is not a condition. Apply it to read-ok.csv and save to /root/bus-triage/i2c-frames.json start_count, stop_count, starts, stops (lists of sample numbers), and scl_high_windows (the number of SCL high intervals confirmed to have gone up and come back down — an interval that stays high to the end has an unknown length, so do not count it).
Judge a condition only when scl[i] and scl[i-1] are both high. You must exclude the moments that overlap with an edge so that you do not mistake a repeated START for data. The grader first checks with short test tables.
Read out the address and ACK
Decode read-ok.csv to the end. Read SDA in the middle of each SCL high interval to collect bits, and for every 9 bits, group the first 8 bits from the high-order bit into a byte, and treat it as an ACK if the ninth bit is 0. Save to /root/bus-triage/i2c-read.json bytes (each entry with value and ack), address (the first byte shifted 1 bit to the right), first_rw, register, second_rw, data (the values after the third byte), repeated_starts (a list of sample numbers), and last_byte_acked. first_rw and second_rw are "read" or "write".
A repeated START does not break the transfer, so reset the flow that counts bytes but continue the transfer. The NACK of the last byte is not a fault but the normal procedure in which the master announces to stop sending.
An address that no one responded to
Decode /opt/fixtures/serialbus/i2c/nack.csv and save to /root/bus-triage/i2c-nack.json start_n, stop_n, byte_count, address, rw, acked (whether the ninth bit was 0), and data_bytes_after_address (the total byte count minus 1).
At the ACK position the master lets go of SDA. That it stayed high means no one pulled it down, and an address error, no power, and a device fault all look the same. Also confirm that the master issues a STOP right away.
An interval the master did not make
Measure all the SCL low interval lengths of /opt/fixtures/serialbus/i2c/stretch.csv and save to /root/bus-triage/i2c-stretch.json scl_low_runs, median_low_us, max_low_us, stretch_us (the maximum minus the median), stretch_start_n (the sample number where the longest interval starts), bytes (the list of decoded bytes), and stop_n.
The reason to use the median is that the mean is pulled by one long interval. If you pin down, by sample number, after which byte the stretched interval lies, you can see when the slave bought time.
The rise uses up the bit time
Measure /opt/fixtures/serialbus/i2c/weak-pullup.csv. Save to /root/bus-triage/i2c-pullup.json sda_rise_time_us and scl_rise_time_us (both from 10 % to 90 %), scl_high_windows (same definition as step 2), sda_peak_v (the maximum of the maximum SDA values within each SCL high interval), ones_at_1v65 and ones_at_vih (the number of intervals whose maximum is at or above 1.65 V and VIH respectively), vih_v (0.7 times 3.30 rounded to two decimal places), and stops_at_1v65 and stops_at_vih (the number of STOP conditions detected when only the SDA decision threshold is changed).
If you compare the rise times of SDA and SCL, you can immediately see which line's pull-up is weak. If the STOP count differs by threshold, it means the slow rise spills into the SCL high interval and looks like a condition.
Which of the four combinations
Implement decode(cs_v, sclk_v, mosi_v, miso_v, cpol, cpha, threshold_v=1.65) in /root/bus-triage/spi.py. While CS is low, sample at the edge going to (1 minus cpol) if cpha is 0, and at the edge going to cpol if it is 1. Return {"edges": the number of sampled edges, "mosi": a list of bytes, "miso": a list of bytes, "leftover_bits": the remainder of the edge count divided by 8} and group bits from the high-order bit. Judge /opt/fixtures/serialbus/spi/mode0.csv and mode3.csv each, and save to /root/bus-triage/spi-mode.json two blocks, mode0 and mode3, each with sclk_idle, cpol, cpha, mode, sampling_edges, mosi_changes_on, mosi, and miso.
cpol is the clock level of the first sample. Determine cpha from the direction of the clock edge just before (within 8 samples of) the time MOSI changes — that edge is not the reading side. If that direction is the same as the idle level, cpha is 0, and if different, 1. mode is cpol times 2 plus cpha.
If the select line is late, the count tells you first
Decode /opt/fixtures/serialbus/spi/cs-late.csv as mode 0 and save to /root/bus-triage/spi-cs.json cs_low_start and cs_low_end (the sample where CS fell and the sample where it rose again), sclk_edges_total (the number of clock edges in the whole capture), sampling_edges_total (half of that), sampling_edges_in_cs (the number of edges sampled while CS is low), leftover_bits, and bytes (the list of MOSI bytes).
The count becomes abnormal before the value. If there is an edge that dropped outside CS, the multiple of 8 breaks and leftover bits arise. Bytes grouped in that state are all shifted by one slot.
Sort out three captures with no names
Judge capture-1.csv, capture-2.csv, and capture-3.csv in /opt/fixtures/serialbus/unknown/ and save to /root/bus-triage/triage.json three blocks, capture_1, capture_2, and capture_3. All three blocks contain channels (the number of channels), bus (one of "uart", "i2c", and "spi"), and fault (one of "polarity-inverted", "sda-stuck-low", "miso-idle", "baud-mismatch", "address-nack", "clock-stretch", "weak-pullup", "cs-timing", and "mode-mismatch"). In addition, capture_1 contains idle_v, baud, and text, capture_2 contains start_count, stop_count, and longest_sda_low_us, and capture_3 contains miso_edges_in_cs, miso_bytes, and mosi_bytes.
The number of channels and the idle level alone nearly settle the bus. If it is one line resting low, try reading with every bit flipped. If it is two lines with no STOP, see what is blocking the STOP. A line in four lines that never changes during the whole transfer is evidence in itself.