eRacks Ships Benchmarked Dual Intel Arc Pro B70 AI Servers and Prices AI Provisioning at $1,495
The Californer/10355599

Trending...
Measured on the company's own bench: 54 tokens per second on Qwen3-14B, 26 on Qwen3.6-27B from a single card - and the provisioning behind those numbers is now a standard priced offering: $1,495, or $2,495 with a private RAG stack, included on flags

CAMPBELL, Calif. - Californer -- eRacks Open Source Systems today published benchmark results measured on a dual Intel Arc Pro B70 server during its pre-ship provisioning pass, and announced that the provisioning work behind those numbers is now a standard priced offering across its AI server line.

The measured numbers, from the company's bench this week: Qwen3-14B generating 54 tokens per second, and the larger Qwen3.6-27B holding a sustained 26 tokens per second on a single B70 - faster than most people read. The serving stack is fully open source: llama.cpp's official Intel build in rootless Podman containers, exposing the industry-standard OpenAI-compatible API, with models resident entirely in GPU memory.

More on The Californer
The hardware is the point. The Arc Pro B70 carries 32GB of VRAM (the GPU's onboard memory, the hard limit on what models fit) per card, so a two-card server fields 64GB of GPU memory for less than the list price of a single 96GB flagship datacenter card. For private AI, that ratio of memory to dollars is the value play of 2026.

"Benchmarks on a spec sheet are marketing. Benchmarks on your machine are engineering," said Joseph Wolff, founder and CTO of eRacks Systems. "Getting these cards to production took three fixes you will not find in any manual - GPU power management that puts cards to sleep permanently, container networking that resets every connection while the server looks healthy. We solved them on the bench, and every AI server we ship now leaves with its own measured numbers and the rebuild notes in the customer's hands."

That work is now a named product: eRacks AI Provisioning & Setup covers burn-in, GPU bring-up with every fix applied, deployment and benchmarking of the customer's chosen models on the customer's actual hardware, and full rebuild documentation - $1,495, or $2,495 including a private RAG stack (retrieval-augmented generation: chat plus a vector database answering from the customer's own documents, fully offline). It is included at no charge on flagship orders.

More on The Californer
The line ships configured to order at live prices: eRacks/AIDAN with one B70 from $13,895, eRacks/AINSLEY with two B70s and 64GB of GPU memory - the configuration class benchmarked above - from $21,395, the full AI server line from $7,695, and the 8-GPU eRacks/HIGHLANDER flagship from $154,995. Configuration and the company's no-signup rent-versus-own calculator: https://eracks.com/products/ai-rackmount-servers/ and https://eracks.com/tco/

Contact
Joseph Wolff, eRacks Open Source Systems
***@eracks.com


Source: eRacks Open Source Systems
Filed Under: Computers

Show All News | Disclaimer | Report Violation

0 Comments

Latest on The Californer