TT Lab
Get started
Learn Learning paths Courses

The Sensor Is Fine. The Board Cannot Understand It

I Do Not Even Know Which Bus This File Is

Continue in TT Lab

Goal

Decode four I2C captures and three SPI captures yourself, and at the end receive three captures with no names attached, sort out which bus each is, and report the faults with numeric evidence attached.

Why it matters

The one line "the device is not responding" that a driver puts out lumps together four different causes. The address is wrong, the slave is buying time, the pull-up is weak so the rise uses up the whole bit time, or the bus is locked up entirely. These four look completely different on a waveform, and the people who fix them differ too — firmware, the driver's recovery procedure, and hardware. If you sort them out from the waveform, one line of the report points to one person in charge. SPI, conversely, comes with its clock, so there is no speed problem, and instead which edge you read on and when the select line falls create quiet faults.

Steps

  1. Read a two-line capture — Measure the sample rate, the rise time, and the distribution of the SCL low intervals.
  2. Find START and STOP — Count only SDA changes while SCL is high as conditions.
  3. Read out the address and ACK — Recover the bytes and the ACKs from groups of 9 bits.
  4. An address that no one responded to — Confirm what a NACK is actually the absence of.
  5. An interval the master did not make — Find clock stretching from the length distribution.
  6. The rise uses up the bit time — Prove a weak pull-up with the rise time and two thresholds.
  7. Which of the four combinations — Determine the mode from the idle level and the edge where the data changes.
  8. If the select line is late, the count tells you first — See the multiple of 8 break.
  9. Sort out three captures with no names — Report the bus and the fault together with evidence.

Notes

Read a two-line capture

Create a working folder with mkdir -p /root/bus-triage, read /opt/fixtures/serialbus/i2c/read-ok.csv, and save to /root/bus-triage/levels.json sample_rate_hz, sample_count, duration_us, channels (the list of channel names in the header), scl_rise_time_samples, sda_rise_time_samples, scl_low_runs, and scl_low_median_us. The rise time is the number of samples of the last complete rise in the capture file (from 0.33 V or below to 2.97 V or above), and an SCL low interval runs from the sample where it went down to the sample where it came back up.

0.33 V and 2.97 V are 10 % and 90 % of 3.30 V. A low interval is between a falling edge and the next rising edge, and an interval that ends still low at the very end has an unknown length, so do not count it.

Find START and STOP

Implement conditions(scl_v, sda_v, threshold_v=1.65) in /root/bus-triage/i2c.py. While SCL is high for two consecutive samples, if SDA falls, put ("start", i) in the list, and if it rises, ("stop", i), and return them in time order. An SDA change while SCL is low is not a condition. Apply it to read-ok.csv and save to /root/bus-triage/i2c-frames.json start_count, stop_count, starts, stops (lists of sample numbers), and scl_high_windows (the number of SCL high intervals confirmed to have gone up and come back down — an interval that stays high to the end has an unknown length, so do not count it).

Judge a condition only when scl[i] and scl[i-1] are both high. You must exclude the moments that overlap with an edge so that you do not mistake a repeated START for data. The grader first checks with short test tables.

Read out the address and ACK

Decode read-ok.csv to the end. Read SDA in the middle of each SCL high interval to collect bits, and for every 9 bits, group the first 8 bits from the high-order bit into a byte, and treat it as an ACK if the ninth bit is 0. Save to /root/bus-triage/i2c-read.json bytes (each entry with value and ack), address (the first byte shifted 1 bit to the right), first_rw, register, second_rw, data (the values after the third byte), repeated_starts (a list of sample numbers), and last_byte_acked. first_rw and second_rw are "read" or "write".

A repeated START does not break the transfer, so reset the flow that counts bytes but continue the transfer. The NACK of the last byte is not a fault but the normal procedure in which the master announces to stop sending.

An address that no one responded to

Decode /opt/fixtures/serialbus/i2c/nack.csv and save to /root/bus-triage/i2c-nack.json start_n, stop_n, byte_count, address, rw, acked (whether the ninth bit was 0), and data_bytes_after_address (the total byte count minus 1).

At the ACK position the master lets go of SDA. That it stayed high means no one pulled it down, and an address error, no power, and a device fault all look the same. Also confirm that the master issues a STOP right away.

An interval the master did not make

Measure all the SCL low interval lengths of /opt/fixtures/serialbus/i2c/stretch.csv and save to /root/bus-triage/i2c-stretch.json scl_low_runs, median_low_us, max_low_us, stretch_us (the maximum minus the median), stretch_start_n (the sample number where the longest interval starts), bytes (the list of decoded bytes), and stop_n.

The reason to use the median is that the mean is pulled by one long interval. If you pin down, by sample number, after which byte the stretched interval lies, you can see when the slave bought time.

The rise uses up the bit time

Measure /opt/fixtures/serialbus/i2c/weak-pullup.csv. Save to /root/bus-triage/i2c-pullup.json sda_rise_time_us and scl_rise_time_us (both from 10 % to 90 %), scl_high_windows (same definition as step 2), sda_peak_v (the maximum of the maximum SDA values within each SCL high interval), ones_at_1v65 and ones_at_vih (the number of intervals whose maximum is at or above 1.65 V and VIH respectively), vih_v (0.7 times 3.30 rounded to two decimal places), and stops_at_1v65 and stops_at_vih (the number of STOP conditions detected when only the SDA decision threshold is changed).

If you compare the rise times of SDA and SCL, you can immediately see which line's pull-up is weak. If the STOP count differs by threshold, it means the slow rise spills into the SCL high interval and looks like a condition.

Which of the four combinations

Implement decode(cs_v, sclk_v, mosi_v, miso_v, cpol, cpha, threshold_v=1.65) in /root/bus-triage/spi.py. While CS is low, sample at the edge going to (1 minus cpol) if cpha is 0, and at the edge going to cpol if it is 1. Return {"edges": the number of sampled edges, "mosi": a list of bytes, "miso": a list of bytes, "leftover_bits": the remainder of the edge count divided by 8} and group bits from the high-order bit. Judge /opt/fixtures/serialbus/spi/mode0.csv and mode3.csv each, and save to /root/bus-triage/spi-mode.json two blocks, mode0 and mode3, each with sclk_idle, cpol, cpha, mode, sampling_edges, mosi_changes_on, mosi, and miso.

cpol is the clock level of the first sample. Determine cpha from the direction of the clock edge just before (within 8 samples of) the time MOSI changes — that edge is not the reading side. If that direction is the same as the idle level, cpha is 0, and if different, 1. mode is cpol times 2 plus cpha.

If the select line is late, the count tells you first

Decode /opt/fixtures/serialbus/spi/cs-late.csv as mode 0 and save to /root/bus-triage/spi-cs.json cs_low_start and cs_low_end (the sample where CS fell and the sample where it rose again), sclk_edges_total (the number of clock edges in the whole capture), sampling_edges_total (half of that), sampling_edges_in_cs (the number of edges sampled while CS is low), leftover_bits, and bytes (the list of MOSI bytes).

The count becomes abnormal before the value. If there is an edge that dropped outside CS, the multiple of 8 breaks and leftover bits arise. Bytes grouped in that state are all shifted by one slot.

Sort out three captures with no names

Judge capture-1.csv, capture-2.csv, and capture-3.csv in /opt/fixtures/serialbus/unknown/ and save to /root/bus-triage/triage.json three blocks, capture_1, capture_2, and capture_3. All three blocks contain channels (the number of channels), bus (one of "uart", "i2c", and "spi"), and fault (one of "polarity-inverted", "sda-stuck-low", "miso-idle", "baud-mismatch", "address-nack", "clock-stretch", "weak-pullup", "cs-timing", and "mode-mismatch"). In addition, capture_1 contains idle_v, baud, and text, capture_2 contains start_count, stop_count, and longest_sda_low_us, and capture_3 contains miso_edges_in_cs, miso_bytes, and mosi_bytes.

The number of channels and the idle level alone nearly settle the bus. If it is one line resting low, try reading with every bit flipped. If it is two lines with no STOP, see what is blocking the STOP. A line in four lines that never changes during the whole transfer is evidence in itself.