<?xml version="1.0" encoding="utf-8"?>
<?xml-stylesheet type="text/xsl" href="../assets/xml/rss.xsl" media="all"?><rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>TinyComputers.io (Posts about engineering margin)</title><link>https://tinycomputers.io/</link><description></description><atom:link href="https://tinycomputers.io/categories/engineering-margin.xml" rel="self" type="application/rss+xml"></atom:link><language>en</language><copyright>Original site content © 2022–2026 Tiny Machines Workshop, LLC, except where otherwise noted. Some rights reserved.</copyright><lastBuildDate>Fri, 04 Sep 2026 15:38:40 GMT</lastBuildDate><generator>Nikola (getnikola.com)</generator><docs>http://blogs.law.harvard.edu/tech/rss</docs><item><title>The Same Bad Luck, Quietly: Why a Fleet of Cheap Boards Feels Flakier Than One Old Xeon</title><link>https://tinycomputers.io/posts/the-same-bad-luck-quietly.html?utm_source=feed&amp;utm_medium=rss&amp;utm_campaign=rss</link><dc:creator>A.C. Jokela</dc:creator><description>&lt;div class="audio-widget"&gt;
&lt;div class="audio-widget-header"&gt;
&lt;span class="audio-widget-icon"&gt;🎧&lt;/span&gt;
&lt;span class="audio-widget-label"&gt;Listen to this article&lt;/span&gt;
&lt;/div&gt;
&lt;audio controls preload="metadata"&gt;
&lt;source src="https://tinycomputers.io/the-same-bad-luck-quietly_tts.mp3" type="audio/mpeg"&gt;
&lt;/source&gt;&lt;/audio&gt;
&lt;div class="audio-widget-footer"&gt;29 min · AI-generated narration&lt;/div&gt;
&lt;/div&gt;

&lt;p&gt;Everything I run lives in one small wall-mounted 19-inch rack in the basement. A Xeon E5-1620 on a Supermicro X9SRL-F board, 128 gigabytes of RAM, does the unglamorous batch work down at the bottom. The fleet lives directly on top of it, and on the forced-air ducting above that: Orange Pi, NanoPi, Raspberry Pi, a Banana Pi or two, a RISC-V board that identifies as a Mars. And here is an observation I can't shake, stated as unscientifically as I experience it: the boards are flaky and the Xeon is not. The boards hang, drop their USB, corrupt a card, come back wrong after a power blip. The Xeon just runs. It has been just running for so long that I've stopped thinking of "running" as something it does, in the way you stop thinking of the furnace as something that works.&lt;/p&gt;
&lt;p&gt;The obvious response is that I'm comparing $80 hobby boards to commercial server equipment, and of course the server wins. True. But "you get what you pay for" is a shrug, not an explanation. The interesting question is what, precisely, the extra money bought, because the one thing it demonstrably didn't buy is speed. &lt;a href="https://tinycomputers.io/posts/freebsd-on-a-2011-macbook-pro.html"&gt;Last month I ran a 2011 MacBook Pro through the same Rust compile benchmark as the SBC fleet&lt;/a&gt;, and once you normalize per core and per gigahertz, a Sandy Bridge core from 2011 and a Cortex-A76 from 2023 do the same work to within a fraction of a percent. Meanwhile &lt;a href="https://tinycomputers.io/posts/friendlyelec-nanopc-t6n-review.html"&gt;the fastest machine on my bench is a $150 FriendlyElec board&lt;/a&gt;. And the Xeon holding up the rack is Sandy Bridge too, a 2012 part wearing a 2018 BIOS, the same silicon generation as that laptop. Whatever makes it dependable, it isn't that it's newer than the boards sitting on it. It is fourteen years old. Performance is not what went missing at the bottom of the market. Something else did.&lt;/p&gt;
&lt;p&gt;Here's the thesis: a spec sheet tells you what a board does. It never tells you how far from the edge it does it. The distance from the edge, which engineers call margin, is where reliability actually lives, and it isn't a component you can point to. It's a budget spent invisibly everywhere: guardband on voltage rails, derating on capacitors, extra dies rejected at the fab, validation hours nobody tweets about, ECC bits doing correction you'll never hear about. Because margin is invisible, it is precisely what cost engineering removes first. Removing it is undetectable at purchase time and for months afterward. The bill comes due later, in someone else's rack. Mine, in this case.&lt;/p&gt;
&lt;h3&gt;Margin Is a Budget, Not a Component&lt;/h3&gt;
&lt;p&gt;Take two boards that post identical benchmark numbers. One runs its DRAM controller with generous timing slack, filters every rail with capacitance the schematic reviewer insisted on, and spent six months in a validation lab getting thermally cycled and brownout-tested. The other runs the same silicon at the same clocks with none of that. No review can tell them apart in week one. Over a year, one of them hangs monthly and the other doesn't, and nobody, including the manufacturer, can tell you exactly why, because the difference was never a feature. It was slack, and slack doesn't have a part number.&lt;/p&gt;
&lt;p&gt;This is why the cheap board's spec sheet looks so good. Everything legible survived the cost-down: the core count, the RAM, the 2.5GbE port. Everything illegible was harvested: the hold-up time in the power supply, the guardband in the voltage tables, the burn-in that would have caught the marginal die. At a $35–150 price point every cent of BOM is fought over, and margin is the one line item you can cut with total deniability. Nobody returns a board because its decoupling network is minimal.&lt;/p&gt;
&lt;p&gt;When you buy server gear, an enormous fraction of the price is tests the machine passed that you will never see. That's the product. It's a terrible product to market ("pays for itself in hangs that don't happen"), which is why nobody markets it and everybody who runs both tiers eventually rediscovers it empirically.&lt;/p&gt;
&lt;h3&gt;Power Is the First Casualty&lt;/h3&gt;
&lt;p&gt;If I had to pick the single largest hole in the SBC reliability budget, it's power, at both ends of the cable.&lt;/p&gt;
&lt;p&gt;At the board end, these machines are built around phone- and tablet-grade PMICs. The X-Powers and Rockchip parts in that rack trace straight back to devices with batteries, which are the world's most forgiving power supplies. At the wall end, they're fed by whatever brick was in the drawer, over connectors chosen for cost. The Raspberry Pi 4 famously shipped with a USB-C detection circuit that caused e-marked cables to identify it as an audio accessory and refuse to power it at all. A small masterpiece of the genre, because it means even doing the &lt;em&gt;right&lt;/em&gt; thing (buying a good cable) could produce a mystery failure.&lt;/p&gt;
&lt;p&gt;The deeper problem is that undervoltage never announces itself as undervoltage. A Pi flags trouble when the rail sags below about 4.63 volts, and a marginal supply will dip past that for a few milliseconds exactly when all the cores wake up and a peripheral draws inrush at once. What you observe is displaced from the cause: USB devices dropping, an Ethernet PHY resetting, a silent hang, an SD card that's suddenly unhappy. Nothing says "power." There's a firmware register that keeps sticky bits recording whether undervoltage has &lt;em&gt;ever&lt;/em&gt; occurred since boot (&lt;code&gt;vcgencmd get_throttled&lt;/code&gt;), and I'd wager most fleets have never had it checked once, which means the most common fault in the tier is also the least diagnosed. The Pi does try to warn you: it draws a small lightning-bolt icon in the corner of the display when the rail sags. On a headless fleet, that's the one instrument the platform has, pointed at a monitor nobody attached.&lt;/p&gt;
&lt;p&gt;Now look at the machine they're sitting on. The Xeon's CPUs sit behind multi-phase VRMs with monitoring on every phase. Its power supply has real power-factor correction and hold-up time: stored energy, measured in milliseconds by spec, to coast through a mains flicker that would drop a wall wart flat. Machines of this class routinely run redundant supplies on conditioned circuits, and the platform logs every excursion where a human can read it later. The boards stacked on top of it experience the same utility power, the same flickers, the same transients. Not similar power: the same power, off the same circuit, from inches away. One machine rides through it and writes a log line. The others reboot and keep no diary.&lt;/p&gt;
&lt;h3&gt;Camera Media, Server Duty&lt;/h3&gt;
&lt;p&gt;The second-largest hole is storage, and it's really a category error we've all agreed to stop noticing. The SD card was designed for cameras: large sequential writes, occasionally, on a device that's off most of the time. We put them under operating systems that write small, constantly, forever (logs, databases, journal commits), and act surprised at the results.&lt;/p&gt;
&lt;p&gt;A consumer SD card has a minimal flash translation layer, minimal wear leveling, and no power-loss protection whatsoever. That last one is the killer. When power drops mid-write, you don't just risk your filesystem; you risk the card's own internal bookkeeping, the FTL's map of which physical blocks hold which logical ones. Corrupt the filesystem and &lt;code&gt;fsck&lt;/code&gt; can help you. Corrupt the map and no tool on your side of the interface can, because the card has forgotten where your data &lt;em&gt;is&lt;/em&gt;. This is how a routine brownout becomes "the card died," and why the failure rate correlates with exactly the boards that also have marginal power. Even the rating system conspires: the A1/A2 marks people shop by certify random-I/O &lt;em&gt;speed&lt;/em&gt;, not endurance or power-fail behavior. Which is to say, neither of the two properties a root filesystem actually needs.&lt;/p&gt;
&lt;p&gt;An enterprise SSD is built on the opposite assumption. There's a bank of capacitors on the PCB whose only job is to flush the DRAM cache when power fails, and the firmware was validated on rigs that yank the plug thousands of times looking for the one sequence that loses data. Above the drive, the server's whole storage stack assumes failure too: journals, checksums, RAID, SMART telemetry. I've been circling this fix for years without quite naming the principle: software RAID5 on a &lt;a href="https://tinycomputers.io/posts/pine64-rockpro64-sata-software-raid5.html"&gt;ROCKPro64 back in 2023&lt;/a&gt;, NVMe on every board that offers a lane today. Getting the root filesystem off SD is the single cheapest reliability upgrade in this hobby, and eMMC or an industrial pSLC card is a respectable middle ground when it isn't an option.&lt;/p&gt;
&lt;h3&gt;Tablet Silicon, Mainframe Silicon&lt;/h3&gt;
&lt;p&gt;Underneath both of those problems is a question of lineage, and this is where the essay earns its philosophy tag.&lt;/p&gt;
&lt;p&gt;The SoCs in that rack descend from tablets and set-top boxes. Allwinner and Rockchip built their businesses shipping silicon into devices that sit in a human's hands or under a human's television, devices whose fault-recovery mechanism is an annoyed person with a power button. That design center explains everything downstream. One DVFS voltage table ships for every die ever fabbed, because per-chip binning costs money and a tablet that crashes occasionally still sells. Firmware is validated to the standard of "the demo image boots." If your particular die happens to be marginal at the top operating point, you will never be told; you'll just own the board that hangs every few weeks, and you'll call it flaky, and the adjective will feel like an explanation.&lt;/p&gt;
&lt;p&gt;Server silicon descends from the opposite tradition. RAS (reliability, availability, serviceability) is literally IBM mainframe vocabulary, and it encodes an inverted assumption: no human is watching, so the machine must notice, correct, log, and contain its own faults. That's machine-check architecture. That's ECC not just on DRAM but on caches and interconnects. That's a BMC, an entire second computer whose only job is to watch the first computer and take notes. The fault-recovery subsystem of a server is &lt;em&gt;the server&lt;/em&gt;. The fault-recovery subsystem of an SBC is you.&lt;/p&gt;
&lt;p&gt;I wrote earlier this year about &lt;a href="https://tinycomputers.io/posts/why-some-chips-last-40-years.html"&gt;why chips like the Z80 survive four decades&lt;/a&gt;, and one of the answers was that they're simple enough to be completely known: every erratum found, documented, and worked around a generation ago. A modern eight-core SoC with an NPU bolted on has something close to mainframe complexity with a tablet's validation budget, which is the least comfortable quadrant of that chart to live in. Complexity you can't afford to verify is just latent failure with good benchmarks.&lt;/p&gt;
&lt;p&gt;There's a Heideggerian footnote here that I can't resist, given &lt;a href="https://tinycomputers.io/categories/heidegger.html"&gt;the running series&lt;/a&gt;. Heidegger's famous observation about tools is that equipment &lt;em&gt;withdraws&lt;/em&gt; when it works. The hammer disappears into the hammering, and only announces itself as an object when it breaks. By that standard the Xeon achieved ideal toolhood years ago; I forget it exists, which is the highest compliment infrastructure can receive. The fleet, meanwhile, keeps stepping forward out of the workflow to be &lt;em&gt;present&lt;/em&gt;: conspicuous, obtrusive, demanding to be understood rather than used. A flaky computer is one that keeps insisting on being a thing.&lt;/p&gt;
&lt;h3&gt;The Xeon Isn't Error-Free. It's Error-Correcting.&lt;/h3&gt;
&lt;p&gt;Here's the part that reframed the whole question for me: the solid machine and the flaky machines are probably having similar amounts of bad luck. They differ in what happens next.&lt;/p&gt;
&lt;p&gt;Google's 2009 fleet study (Schroeder, Pinheiro, and Weber, still the reference on this) found that correctable DRAM errors are not rare events but background weather: on the order of a third of machines and more than 8% of DIMMs took at least one correctable error per year, with affected modules often logging thousands. On a platform with ECC, every one of those events is a corrected word and an incremented counter that nobody reads. On a board with no ECC, no machine-check reporting, and no BMC, the &lt;em&gt;same physical event&lt;/em&gt;, one flipped bit from a marginal cell or a stray particle, is a mystery. If it lands in the wrong place, it's a hang with no witnesses, no log line, no autopsy.&lt;/p&gt;
&lt;p&gt;And this generalizes into what I think is the real answer to the flakiness question. On the fleet, every distinct fault class (a voltage sag, a flipped bit, a driver race, a thermal excursion, a dying card) collapses into one undifferentiated symptom: &lt;em&gt;it got weird&lt;/em&gt;. Without witnesses you can't do differential diagnosis, and without diagnosis you can't accumulate knowledge, so the whole tier gets described with an adjective instead of a distribution. "Flaky" is what fault data looks like when nothing is recording it.&lt;/p&gt;
&lt;p&gt;Then add plain arithmetic. There are nine-ish boards in the rack and one server holding them up. At &lt;em&gt;equal&lt;/em&gt; per-unit failure rates the fleet generates nine times the incidents, and the rates are not equal, for every reason above. Multiply an elevated rate by a fleet-sized N, subtract all telemetry, and the subjective experience is exactly what I have: a big machine that never does anything, and a stack of boards on top of it that's always doing &lt;em&gt;something&lt;/em&gt;. Survivorship of attention finishes the job: the Xeon's corrected errors never interrupt me, so as far as experience is concerned, they never happened.&lt;/p&gt;
&lt;h3&gt;Software Is Margin, Too&lt;/h3&gt;
&lt;p&gt;Everything so far was hardware, but the same asymmetry runs through the software, because software maturity is just margin accumulated in public.&lt;/p&gt;
&lt;p&gt;The Xeon runs a kernel hardened by millions of identical deployments over two decades; every erratum in that core has had a microcode or kernel workaround shipping for ten years. Each SBC, meanwhile, lives on a vendor BSP kernel pinned to some ancient LTS, a hand-maintained device tree, and an Ethernet PHY whose RGMII timing quirk gets fixed only when some hobbyist bisects the kernel on a weekend. The entire linux-sunxi community exists because Allwinner's own code needed a volunteer rescue mission; Armbian is best understood as the distributed QA department these vendors never hired. I have my own receipts scattered across this site: &lt;a href="https://tinycomputers.io/posts/building-a-kernel-and-disk-image-for-the-radxa-cm3.html"&gt;hand-building a kernel and disk image just to make a Radxa CM3 usable&lt;/a&gt;, &lt;a href="https://tinycomputers.io/posts/four-partitions-and-a-borrowed-bootloader-netbsd-on-the-milk-v-mars.html"&gt;borrowing a vendor bootloader to get NetBSD onto the Milk-V Mars&lt;/a&gt;, and &lt;a href="https://tinycomputers.io/posts/horizon-robotics-x3-cm-review.html"&gt;the Horizon X3, a board whose software situation earned the phrase "cautionary tale" in the title&lt;/a&gt;. None of those machines was unreliable silicon, exactly. They were unfinished &lt;em&gt;systems&lt;/em&gt;, shipped at the moment the demo booted, with the last 20% of engineering left as an exercise for the buyer.&lt;/p&gt;
&lt;h3&gt;Buying It Back&lt;/h3&gt;
&lt;p&gt;So the fleet is missing margin at every layer. Here is the redemptive turn, and the reason I don't consider any of this a complaint: most of that margin can be bought back, and the project of buying it back is philosophically respectable. It is, in fact, the founding bet of the modern datacenter. Early Google famously chose cheap, failure-prone commodity hardware and moved reliability up into software, treating machines as cattle rather than pets. Running a stack of $80 boards well is that same bet at 1:1000 scale.&lt;/p&gt;
&lt;p&gt;Concretely, four moves recover most of it. First, take power seriously: a real 5A supply per board instead of drawer archaeology, and actually read &lt;code&gt;get_throttled&lt;/code&gt;. The fleet will confess to brownouts you never noticed. A small UPS in front of the whole rack costs less than one board and deletes the mid-write brownout, the nastiest failure in the tier, at its source. Second, get the OS off raw SD: NVMe or eMMC where possible, and where it isn't, a read-only root with an overlay filesystem, so a power cut has nothing in flight to destroy; &lt;code&gt;raspi-config&lt;/code&gt; will set this up on a Pi, and Armbian can do the same. Third, arm the watchdog. Nearly every SoC in that rack ships a hardware watchdog timer that sits idle by default; one line of systemd configuration (&lt;code&gt;RuntimeWatchdogSec=10&lt;/code&gt;), and a silent hang stops being a Saturday trip down to the basement and becomes a fifteen-second self-reboot with a log entry. Fourth, add witnesses: a metrics exporter and a scrape target turn "flaky" from an adjective back into a distribution: &lt;em&gt;which&lt;/em&gt; board, at &lt;em&gt;what&lt;/em&gt; temperature, correlated with &lt;em&gt;what&lt;/em&gt; workload, how often. Monitoring is retrofitted telemetry; it's the poor man's BMC.&lt;/p&gt;
&lt;p&gt;And then there's the software-redundancy move, which I backed into this summer when &lt;a href="https://tinycomputers.io/posts/nothing-was-using-the-cores-three-idle-arm64-boards-one-ha-k3s-cluster.html"&gt;three idle boards in that rack became a three-node HA k3s cluster&lt;/a&gt;. The reliability effect surprised me more than the utilization effect. A node hang stopped being an outage and became a log line and a rescheduled pod, the exact transformation ECC performs on a bit flip, implemented three layers up. Better: faults finally got &lt;em&gt;names&lt;/em&gt;. The big.LITTLE crash loop that surfaced during that build would, on a lone board, have been one more entry in the ledger of weirdness; inside a cluster with health checks and logs, it was observable long enough to be understood.&lt;/p&gt;
&lt;p&gt;What can't be bought back is worth naming honestly: ECC. With rare industrial exceptions, the memory on these boards is soldered down and the controllers were never taught to correct it. That last piece of margin was removed at the fab, and no amount of software puts it back. You can only surround it (checksum the data, replicate the state, watchdog the hang) and accept that some small residue of mystery is the price of the tier.&lt;/p&gt;
&lt;h3&gt;The Same Bad Luck, Quietly&lt;/h3&gt;
&lt;p&gt;Which brings the philosophical question back around. Server hardware is engineered on the assumption that things go wrong: power sags, bits flip, drives lie, and the machine's job is to notice, correct, and keep a diary. Consumer silicon is engineered on the assumption that things mostly won't, and that a human will be standing nearby when they do. The Xeon at the bottom of the rack isn't luckier than the boards stacked on top of it. It is having the same bad luck, quietly, absorbing the same flickers and flipped bits I never hear about, because someone spent money, decades ago and again at the factory, on making bad luck boring.&lt;/p&gt;
&lt;p&gt;That's the real answer to my own question, and it's also why I've come to see the fleet's flakiness as something closer to candor than defect. The boards show me every fault the big machine hides, and in exchange they've taught me the entire discipline the big machine lets me skip: power budgets, storage physics, watchdogs, the value of a witness. How much of the missing margin can be restored from above is, I've realized, the central question of this whole tier of computing, and probably of this site. The answer so far: most of it. Not all of it. The remainder is either the reason the Xeon stays plugged in, or the tuition for everything the fleet keeps teaching. Some weeks it's both.&lt;/p&gt;</description><category>ecc</category><category>engineering margin</category><category>heidegger</category><category>home lab</category><category>philosophy</category><category>power delivery</category><category>reliability</category><category>sd cards</category><category>single board computers</category><category>watchdog</category><guid>https://tinycomputers.io/posts/the-same-bad-luck-quietly.html</guid><pubDate>Fri, 04 Sep 2026 14:00:00 GMT</pubDate></item></channel></rss>