Campaigns ·

Error Independence: does different hardware buy a second opinion?

Measured 150 items, four model configurations

The same model on different silicon shares its errors (phi +0.447 on contested items); models from different families do not (mean phi −0.102).

Registration and verificationGround truth is generated with each item, so no model grades another.

The question

A common resilience design runs the same model a second time, on different accelerators or in a second region, and treats the second answer as a check on the first. Civil aviation learned long ago that two identical computers fail the same way at the same moment. This study asked whether a second copy of a model, on different hardware, makes different mistakes.

What was registered

Ground truth is generated with each item, so the correct answer is known before any model sees it and no model grades another.

What was measured

Four model configurations scored the same 150 invoice-extraction items. The key comparison ran qwen3.8 twice, once as NVFP4 on SGLang on an NVIDIA node and once as Q4_K_M on ollama on an AMD card: identical weights, different quantisation, silicon and serving stack. For each pair the study computed the phi coefficient on error indicators, near zero when two models go wrong on different items and near one when they go wrong together.

103 of the 150 items were unanimous, 54 that every model got right and 49 that every model got wrong. Those say how hard an item is and nothing about shared weakness, so the headline uses the 47 items where the models disagreed.

What it found

Before conditioning, every pair looked correlated, from 0.556 to 0.827. On the contested items, the pair with identical weights on different hardware stayed at +0.447, the highest in the matrix. The five pairs from different model families averaged −0.102, and qwen3.8 NVFP4 against gemma3 sat at −0.469; gemma3 against qwen3.6 was −0.105. Changing the hardware, the quantisation and the serving stack left the correlation where it was. Changing the model family moved it through zero.

Headline numbers, as published
Error correlation (phi), identical weights on different silicon, contested items only+0.447
Mean error correlation (phi) across five cross-family pairs, contested items only (range −0.469 to +0.280)−0.102
Error correlation (phi) across all pairs before conditioning on item difficulty0.556 to 0.827
−0.50.00.51.0open ring: all 150 items · dot: 47 contested itemsSame weights, two machinesSame weights, two machines, all 150 items: 0.827Same weights, two machines, 47 contested items: 0.4470.447qwen3.8 Q4 + qwen3.6qwen3.8 Q4 + qwen3.6, all 150 items: 0.810qwen3.8 Q4 + qwen3.6, 47 contested items: 0.2800.280qwen3.8 NVFP4 + qwen3.6qwen3.8 NVFP4 + qwen3.6, all 150 items: 0.729qwen3.8 NVFP4 + qwen3.6, 47 contested items: 0.0130.013gemma3 + qwen3.6gemma3 + qwen3.6, all 150 items: 0.617gemma3 + qwen3.6, 47 contested items: −0.105−0.105qwen3.8 Q4 + gemma3qwen3.8 Q4 + gemma3, all 150 items: 0.607qwen3.8 Q4 + gemma3, 47 contested items: −0.227−0.227qwen3.8 NVFP4 + gemma3qwen3.8 NVFP4 + gemma3, all 150 items: 0.556qwen3.8 NVFP4 + gemma3, 47 contested items: −0.469−0.469
Pairwise error correlation (phi) on all 150 items (open ring) and on the 47 contested items (filled dot). The highlighted pair runs identical weights on different hardware. Rebuilt from the numbers in When the Second Opinion Shares the Blind Spot.
Show the numbers as a table
PairAll 150 items47 contested items
Same weights, two machines0.8270.447
qwen3.8 Q4 + qwen3.60.8100.280
qwen3.8 NVFP4 + qwen3.60.7290.013
gemma3 + qwen3.60.617−0.105
qwen3.8 Q4 + gemma30.607−0.227
qwen3.8 NVFP4 + gemma30.556−0.469

What it does not claim

  • The same-family reading comes from one pair, the only two systems in the set with identical weights. One pair is an observation.
  • The five cross-family pairs run from −0.469 to +0.280, and the mean sits inside a wide spread without describing it well.
  • One task, chosen because its ground truth cannot be argued with. Open-ended work is untested.

Read the pieces

“Better questions lead to better worlds.”Dinand Tinholt

Follow your curiosity.

Surprise me
Top