🎧 Listen to this article

In September I wrote up the arithmetic for a small RISC-V rental business: two SpacemiT K3 nodes in a 1U at a Miami colo, scaling to ten if anyone cared. I costed the capital, the colocation, the bandwidth and the tariffs. What I could not cost, because I did not own one, was whether the software layer around this silicon was ready to rent to strangers.

So I bought one board. \$889.00 from Firefly, \$45.00 shipping from Hong Kong, and \$241.09 in customs duty that DHL collected on the doorstep. \$1,175.09 to find out.

Three weeks later I have an answer, and it is not the one I expected. The board is considerably better than I thought. The ecosystem around it is considerably worse. Those sound like they cancel out. They do not, because they are answers to two different questions, and the questions resolve in opposite directions.

What the option cost

Line Amount
Board \$889.00
Shipping from Hong Kong \$45.00
Customs duty, collected by DHL on delivery \$241.09
Total \$1,175.09

And what it was an option on:

Line Amount
CSB1-N10SPK3, K3x10 at 32G+128G, BMC3588 8G \$16,419.00
Shipping, quoted, 19.1 to 20 kg band \$240.00
Duty at 27.1%, the rate I was charged in September \$4,449.55
Customs broker, formal entry \$150.00 to \$400.00
Single entry bond \$75.00 to \$275.00
Landed \$21,333.55 to \$21,783.55

Above \$2,500 you leave informal entry behind, and formal entry wants a customs broker and a single entry bond. I have left both as ranges rather than chasing a firm quote, and that is a deliberate choice rather than laziness.

A broker will quote you a number. That number is good for the shipment in front of them, under the schedule in force on the day, and I have no confidence that the schedule in force on the day I might actually order resembles the one in force now. The 27.1 percent I paid in September is itself a single observation from a single shipment. Writing \$21,412.83 would imply a precision that the current state of international trade does not support. A range that brackets the answer is more useful than a point estimate that will be wrong in a more specific way.

So call it \$21,300 to \$21,800 landed, and treat even that as soft at the edges. Against the top of the range, \$1,175.09 was about five and a half percent of the commitment, spent to find out whether the other ninety four percent was worth spending.

The board is better than I expected

I want to lead with this because it is the part I got wrong in the pessimistic direction, and because every number here is one I did not have in September.

I ran twenty four Debian riscv64 guests on it, simultaneously, under KVM, on eight cores. Forty eight virtual CPUs on eight physical ones, six times oversubscribed. Every one booted. None failed. Boot time was ten seconds per guest and stayed ten seconds at twenty four, which is not what I expected from a board this far from the mainstream.

More interesting than the fact that it worked is the shape of how it worked. Aggregate throughput across all guests is linear to three guests, saturates at four, which is exactly eight virtual CPUs on eight cores, and then holds absolutely flat: 5,446 sysbench events per second with a standard deviation of 16 across twenty one saturated levels. A spread of under one percent. Total capacity is fixed and guests divide it. Six fold oversubscription costs nothing in aggregate; it only slices the same pie thinner.

That gives a clean way to state what a tenant gets:

Guests Each gets, as a share of a guest running alone
5 75 percent
7 50 percent
11 33 percent
15 25 percent

Pick the tier you are willing to sell. The board will run twenty four without complaint. Whether a tenant renews at sixteen percent of a dedicated guest is a pricing question, not a hardware one.

Memory never became the constraint. Guests touched 399 MB each against 2,048 MB allocated, leaving 21.8 GB free at twenty four, which extrapolates to roughly seventy eight guests before RAM binds. CPU binds an order of magnitude earlier.

Then I left twelve guests resident and put the board through a duty cycle for a day: idle floor, ramp, sustained, a saturated peak with all twelve working at once, and an evening taper, repeating hourly. Both tenant tiers running, virtual machines and containers, the containers held to a CPU quota the way an operator would actually partition them.

It did not miss. Twenty two cycles, every phase on schedule, twelve of twelve guests answering at every single health check. Zero guest losses. Swap never moved off zero. No new kernel errors inside the run window. The NVMe, a 2TB SanDisk BLACK SN850X, reports zero percent wear and zero media errors after one hundred and seventy power on hours. Input/output wait peaked at 0.3 percent.

The duty cycle has a power signature now, too. Since 7 October the board has been fed through an inline meter, a Shelly Plug US Gen4 smart plug read once a second over its local HTTP API by a FreeBSD machine on the same network, and the harness's own phase log lines up with the wall draw to the minute. Thirty six consecutive cycles, every phase labelled by the harness rather than inferred from the trace:

Phase, per hourly cycle Working guests Container quota Wall power Energy
valley, 6 min 0 of 12 idle 12.9 W 1.3 Wh
ramp, 9 min 1 of 12 0.5 cores 15.1 W 2.3 Wh
midday, 18 min 2 of 12 1.0 core 17.5 W 5.3 Wh
peak, 15 min 12 of 12 1.0 core 19.9 W 4.9 Wh
evening, 12 min 2 of 12 0.5 cores 16.8 W 3.4 Wh

Wall power, current and power factor across the K3 load-harness duty cycle, with harness phase bands and per-phase averages: 12.9 watts at the idle floor rising to 19.9 watts with all twelve guests working

Three things fall out of that table. The first is how narrow the range is: idle floor to every-guest-saturated is seven watts, meaning two thirds of the wall draw is the fixed cost of the board, the brick and twelve resident qemu processes, and the compute itself is the cheap part. The second is the marginal cost of a tenant: going from two working guests to all twelve, 49 to 100 percent of host CPU, adds 2.3 watts, about a quarter of a watt per additional busy guest. Density is nearly free on this board; it is the first tenant that costs. The third is that it does not move. Thirty six cycles, peak reproducible to within a few hundredths of a watt, zero thermal guard trips, and a power factor pinned at 0.64 no matter the phase. Total consumption over the window was 0.41 kilowatt-hours a day, which at fifteen cents a kilowatt-hour is six cents a day to simulate a full rack's weekday.

And it does not throttle. This board exposes no thermal zone, no cpufreq interface and no power domain to mainline Linux, so I cannot read the die temperature, the clock or the wattage. None of them. Rather than guess, I timed a fixed unit of single threaded work on core zero every five minutes for fourteen hours and recorded which phase of the duty cycle each measurement fell in. The idle samples drifted 0.3 percent over 14.1 hours across sixteen measurements. The loaded samples sit at 500 to 680 against an idle 900 to 912. The whole gap is contention between tenants, not heat. The chassis proxies I can read, the network controller die and the NVMe controller, sawtooth with the duty cycle between 45 and 53 degrees and recover fully in every idle phase.

For a rental business this is the good news. The hardware holds up.

The ecosystem is worse than I expected

Now the other half.

Mainline Linux enumerates eight of this chip's sixteen cores. The other eight, the A100 cluster, simply do not exist as far as the kernel is concerned. The vendor kernel sees them but reaches them only through a proprietary hook, and no process can be scheduled across both clusters anyway. You are selling half a chip.

The internal storage does not work. The device tree declares the UFS controller with no voltage regulators and no frequency table, so link training fails and the controller never comes up. On mainline the drive is simply absent. Root lives on NVMe because it has to.

FreeBSD, NetBSD and OpenBSD cannot boot on this board at all, virtualized or not. The K3's M-mode firmware leaves a single bit clear in a control register, mcounteren.CY, and any guest that executes rdcycle traps immediately. All three BSDs do, about a second into boot. This is not a driver gap that patches out. It is one line of firmware configuration that nobody outside SpacemiT can change.

Then there are the container images, which is where the immaturity is easiest to measure. I queried twenty nine images a tenant might plausibly ask for. Nineteen of the twenty seven that resolved publish riscv64 builds. The eight that do not are not obscure: no Node.js at any version, no MySQL, no MariaDB, no MongoDB, no Grafana, no InfluxDB. A LAMP or MEAN stack cannot be assembled from official images on this architecture today. A Postgres and Go stack can, comfortably.

The most telling detail is that support is version gated, and the gate moved recently. redis:7 publishes no riscv64 manifest. redis:8 does. rust:1.83 does not. rust:latest does. So the ordinary defensive habit, pin the major version and do not chase latest, is precisely the habit that breaks on this architecture. A tenant who writes redis:7 because that is what their x86 fleet runs gets a manifest error, concludes the platform does not support Redis, and is wrong by exactly one major version.

One more, because it is the kind of thing you only find by running the machine rather than reading about it. During a fleet boot, one guest in twelve kernel panicked thirty three seconds in, with an instruction page fault on kernel text, which should not be possible. It has not recurred in a day of sustained load. But the panicked guest did not go quiet. It burned 199.7 percent CPU, its entire virtual CPU allocation, indefinitely, while the qemu process looked healthy and the forwarded port still accepted connections. A quarter of the board consumed by a virtual machine doing nothing, invisible from the host. A tenant would have kept paying for it.

Two questions, not one

Here is where the two halves stop cancelling out.

The first question is: should I buy the ten node chassis now? The answer is no, and it was no in September too. Every one of the problems above is a property of the silicon, its firmware, or the software around it. Not one of them is a property of how many boards you own. Buying ten copies does not give you ten chances to discover that the BSDs cannot boot. It gives you one discovery and nine idle nodes. The integration work I have just done, replacing the bootloader, finding out which kernel costs you what, learning the thing cannot be installed the way every other server racks, is done once and does not repeat per node. That is the argument for having done it on one board at 5.5 percent of the price.

The second question is different: is there a business here? And the answer is yes, for a narrow customer, and the reason it exists is the same fact that makes the chassis a bad buy.

The ecosystem gaps keep the large providers out. There is no volume in RISC-V hosting yet, and the tenants who want it want things that do not scale: serial console access, the ability to boot their own kernel, permission to break the machine. That is a bad business for a hyperscaler and a defensible one for somebody small. The immaturity is not an obstacle to the business. It is the moat.

And notice what is actually missing from the image ecosystem. No Node, no MySQL, no Mongo, no Grafana. That is a dealbreaker for a tenant deploying a web stack and completely irrelevant to a tenant who wants Debian, a compiler and root. The gaps are precisely in the software the target customer does not care about.

Who the customer actually is

Not the person who wants to run a LAMP stack. That person should use x86 and will be happier.

The customer is the kernel porter and the firmware developer. But those are two different products, and my own measurements say so.

A kernel porter working in supervisor mode and above is well served by a virtual machine here. When I pass -cpu host the guest inherits the real silicon features, including vector crypto, so what you are renting is access to an actual RVA23 class implementation rather than an emulation of one. Ten second boots matter enormously to someone whose workflow is a reboot loop. Copy on write overlays on a shared golden image give snapshot and revert for free, which is what you need when you are panicking kernels deliberately. All of this works today, measured, on hardware that costs \$1,175.

A firmware developer cannot use a virtual machine at all, and this is the thing I would have got wrong without owning the board. KVM requires the host's OpenSBI to own machine mode; guests run in hypervisor supervisor mode and below. That is exactly why the BSDs cannot boot under KVM here. A tenant inside a virtual machine cannot inspect, fix or experiment with the mcounteren bit that causes the problem, because that bit lives in a privilege level the hypervisor reserves for itself. Firmware work needs bare metal.

Which brings up the operational problem this whole exercise surfaced, and the reason I think the bare metal tier is the defensible one. Recovering this board's firmware requires holding a physical button while power is applied. Not a reset command. Not a power cycle. A human hand, or a machine pretending to be one. That is why my pre-sales questions to the colo were about remote hands rates and response times rather than bandwidth.

So the product that nobody else is selling is the unglamorous plumbing: out of band serial console to every node, remote power control, and an automated path to unbrick a board a tenant has bricked. The Raspberry Pi acting as a baseboard management controller in my September design, productised. It does not scale, which is exactly why it is defensible. Sell bare metal by the hour to the firmware people and virtual machines to the kernel people, on the same hardware.

What the colo actually costs

In September I could not size any of this, because ServerPronto publishes a one amp allowance without saying at what voltage. They answered on 1 October. It is one amp at 120 volts, which is the pessimistic case of the two I flagged.

That single number sorts the two questions faster than anything I measured on the bench.

One amp at 120 volts is 120 watts, and continuous-load practice says size to eighty percent of the breaker, so call it 96 watts usable. My per-node estimate was 25 to 40 watts, and when I first wrote this section it was the one number in the article I had not measured. That is fixed now: the meter says 12.9 watts at the floor and 19.9 watts with every guest working, at the wall, across thirty six duty cycles. I am keeping 25 watts per node as the planning figure anyway, because one board in open air is not ten in 1U and being wrong in that direction costs me money rather than saving it. But it is a conservative planning figure sitting on a measurement now, not a guess.

For the two-node pilot, with a Raspberry Pi 5 acting as the baseboard controller, the metered worst case is 40 watts at the wall for the pair plus a handful for the Pi: call it 50 watts continuous against the 96-watt usable budget. Even the old pessimistic 40 watts per node would fit. The pilot does not need a second amp at any defensible figure, and that question is closed.

For the ten-node chassis it is not close.

Per node Wall draw Circuit needed Additional amps Colo per month
25 W 332 W 4 A 3 \$138.80
30 W 391 W 5 A 4 \$166.75
40 W 508 W 6 A 5 \$194.70

Additional amps are \$27.95 each per month. So the ten-node rig turns a \$54.95 colo bill into something between \$138.80 and \$194.70, every month, forever. The September arithmetic did not have that line in it at all, because I did not know the voltage and did not ask. One honest note in the chassis's favour: at the metered 19.9 watts per node the same arithmetic lands at three amps and about \$111, below the table's best row. Power was never the strongest argument against the chassis. Bandwidth is the one that does not move, and it is next.

Note what this does to the two questions. The hardware I decided not to buy would also have cost two and a half to three and a half times as much to keep plugged in. The hardware I did buy, in the configuration that makes sense, fits inside the base plan.

Then there is bandwidth, which turns out to matter more than I expected.

Thirty megabits is included, with overage billed at \$5 per megabit. That is plenty for one board and nowhere near enough for ten.

Build Tenants Mbps each
Two-node pilot, 5 per node 10 3.00
Two-node pilot, 7 per node 14 2.14
Ten-node, 5 per node 50 0.60
Ten-node, 7 per node 70 0.43

Two to three megabits per tenant is workable for the customer I described, because kernel work is bursty rather than sustained. At the full pipe a tenant pulls the Linux tree in about thirteen minutes, a Debian netinst image in three, a Postgres container in under two. Nobody is streaming video. Ten or fourteen tenants whose bursts do not line up will mostly not notice each other.

Half a megabit per tenant is not workable by any reading, so the ten-node build has to buy pipe. Getting to 100 Mbps means 70 additional megabits at \$5 each, which is \$350 a month on top of everything else.

So the full colo bill, power and bandwidth together, and then the number that actually matters, which is what each tenant has to pay for the thing to break even. Hardware amortised over three years, colo as quoted, setup fee spread across the same term.

Two-node pilot Ten-node chassis
Hardware, landed \$2,750.18 \$21,558.55
Amortised over 36 months \$76.39 \$598.85
Colo, power \$54.95 \$166.75
Colo, bandwidth included \$350.00
Setup, amortised \$2.22 \$2.22
Monthly total \$133.56 \$1,117.82
Tenants at five per node 14 70
Break even, per tenant \$9.54 \$15.97

Five times the hardware does not get you five times cheaper. It gets you a break even that is 1.67 times higher per tenant. Bandwidth does not get cheaper with density, power gets worse, and the chassis costs more per node than a bare board does.

That last point needs a caveat, because it is the one comparison here that is not like for like: the chassis nodes are specced at 32 GB and 128 GB and the price includes a 1U enclosure, a power supply and an RK3588 management controller, none of which my bare board has. The colo figures are directly comparable. The hardware figures are not, and some of that 1.83x per node is buying real things.

But the colo half alone is enough to make the point, and it is the third separate measurement pointing the same direction, after the guest scaling ceiling and the power budget. The chassis I decided not to buy would have been worse per unit, not better.

Two more answers worth recording.

Their tech support runs 24/7 with ten to fifteen minute ticket response, and remote hands is free when a task is simple, takes a few minutes, and is not frequent. Holding a button while power is applied is simple and takes seconds. It is the word frequent that should worry me. The tenant I identified as the defensible customer, the firmware developer on bare metal, is precisely the tenant who bricks a board regularly. A business whose product is unbricking cannot rest on a courtesy extended only to people who rarely need it. That is a conversation to have with them in writing before taking money, not after.

And the setup fee is real: \$79.95 one time, which the third-party directory had right.

What has to change, and how I will know

The uncomfortable part of the argument is that the moat closes as the ecosystem matures, and the same maturity is what would eventually justify the ten node chassis. Both happen at once. So the question is not whether but when, and I would rather publish a test than a prediction.

Here is what I will re-measure, with the tooling that produced the numbers above:

Gate Today How it gets checked
Node, MySQL, MongoDB riscv64 images absent the Docker Hub survey, one command
Mainline enumerates all sixteen cores eight of sixteen the mainline benchmark script
BSDs boot, mcounteren.CY set broken in M-mode firmware the QEMU run scripts
UFS internal storage device tree missing regulators dmesg on a newer kernel
Guest panic rate one at boot, none in 42 h of load since the sustained load harness

The core enumeration gate is the one to watch. Right now the business sells eight cores of a sixteen core chip. If mainline lands the second cluster, capacity roughly doubles on a software update, without buying anything. That single change moves the unit economics more than any hardware decision available to me.

The image gate is already moving in the right direction, and measurably. redis:7 to redis:8, rust:1.83 to rust:latest. That is not a static wall. It is a wall receding at a rate I can put a number on if I sample it twice. So I intend to: same survey, same board, same charts, in six months. The difference will be the article.

My honest guess is that the sweet spot lands somewhere between six and twelve months from now, when enough of those gates have flipped that a tenant can be productive without hand rolling everything, and before enough have flipped that somebody with real money notices. That is a narrow window and I may be wrong about its position. I will not be wrong about whether it opened, because I will have measured it.

What one board still has not told me

A single unit on a desk is not a rack.

I do not know how ten of these behave thermally in 1U. I have one board in open air and it never got close to warm. Density is a different problem, and mainline will not report the temperature even if it were.

I know the power draw now: 12.9 watts at the floor, 19.9 saturated, from a \$30 inline meter, the purchase an earlier version of this section wished for. It should have happened in September, and the answer turned out to be the boring one, which is the best kind. What one metered board cannot tell me is ten sharing 1U of exhaust. One wrinkle it did surface: the power factor is 0.64 at every load level, so the circuit carries about 31 volt-amps at peak to deliver those 19.9 watts. ServerPronto's allowance is denominated in amps, and the measured 0.26-amp peak per node is the figure that allowance should actually be compared against.

I do not know the guest panic rate. One event is an anecdote. A day of load without a repeat makes it look like a boot concurrency artifact rather than something that happens in service, which is reassuring and not the same as knowing.

And I do not know whether anyone will pay for this. Nobody has paid me anything. That remains the least answered question from the September post and the only one that no amount of benchmarking will settle.

The general point

The thing I keep coming back to is not about RISC-V.

I was about to spend somewhere north of \$21,000 on a bet whose critical assumption was that the software layer around a new system on chip was mature enough to rent to strangers. That assumption was wrong. There was no way to check it from a product page, a press release, or a benchmark chart, because the failures are not the kind of thing anybody writes down. They are a bit left clear in a firmware register, a device tree missing a regulator, a container tag that skipped an architecture, a recovery procedure that needs a finger.

Five and a half percent of the commitment, spent up front, bought all of it. And it bought something better than a no, which is a specific list of what would have to change, and the instruments to check.

Caveats. The \$16,419.00 price and \$240.00 shipping are Firefly's own figures as of 27 September 2026, taken from their checkout. The formal entry costs are industry ranges rather than a broker's quote, left that way on purpose: the 27.1 percent duty is one observation from one shipment, tariff schedules move, and a precise landed figure would be precise about conditions nobody can guarantee will hold. The guest scaling and sustained load numbers come from a single board in open air, which is the best case for thermals and the worst case for generalising to a rack. sysbench cpu is an integer benchmark with a small working set, so it flatters shared cache behaviour; a memory bandwidth bound workload would very likely show degradation where mine shows none, and the guests all ran the same thing at the same time, which is the easiest case for a scheduler. The ServerPronto figures are from their ticket #151276 dated 1 October 2026 and are quoted as given. Where their bandwidth answer and their public plan page disagree I have used the answer, which is the more recent and the more generous of the two; anybody ordering should get it restated on the invoice. Per-node wattage is measured at the wall for a single board in open air (12.9 to 19.9 watts by duty phase, Shelly Plug US Gen4, one-second samples across 36 harness cycles), and the 25-to-40-watt planning figures are kept deliberately above the measurement for rack conditions. All of the technical findings are documented with raw output in the benchmark repository.