Table of Contents
Published: 2026-09-18 · Category: RISC-V AI · Reading time: ~5 min · Status: DRAFT
From announcement to shipping silicon
Axelera AI has moved its second-generation Europa AI Processing Unit (AIPU) from announcement to general availability. The architecture was first unveiled in October 2025 with shipments promised for the first half of 2026; the 15 September 2026 launch confirms silicon and accelerator products are now shipping, in Eindhoven-announced form.
This draft covers the shipping configuration and the published throughput numbers. It is a separate milestone from the earlier architecture announcement.
Architecture: AI cores plus RISC-V vector cores
Europa is not a GPU. The die combines two distinct core types:
- 8 × second-generation AI processing cores — the matrix engine, built on Axelera's Digital In-Memory Compute (D-IMC) approach, where SRAM crossbar arrays perform matrix-vector multiplication in place to cut data movement.
- 16 × RISC-V vector processing cores — these handle the non-AI pre- and post-processing around inference: data preparation and output handling.
That split is the reason this part belongs in a RISC-V discussion at all. The vector cores are a real RISC-V cluster doing the work that would otherwise bounce back to the host CPU. Keeping tokenisation, image pre-processing and output post-processing on the accelerator — instead of round-tripping over PCIe — is where a lot of real-world inference latency actually goes.
Alongside: an onboard H.264 / H.265 video decoder, so camera and industrial-inspection workloads do not consume host CPU cycles.
Published specifications
| Item | Value |
|---|---|
| AI cores | 8 × second-generation |
| RISC-V vector cores | 16 |
| Peak compute | 629 TOPS (quoted at INT8 by TechPowerUp) |
| Precision | INT4, INT8, INT16 |
| TDP | 45 W |
| On-chip L2 SRAM | 128 MB |
| Memory interface | 256-bit LPDDR5 |
| Memory bandwidth | 200 GB/s |
| Max memory per chip | up to 64 GB (per configuration) |
| Host interface | PCIe 4.0 ×4 |
| Video | H.264 / H.265 decoder onboard |
| Process | Samsung 5 nm (per independent coverage) |
| Software | Voyager SDK |
Three form factors and the measured numbers
Europa ships as a bare chip for custom boards, and as two cards:
- Axelera Edge 232p — one AIPU on a half-height, half-length (HHHL) PCIe card. Up to 629 TOPS. Published result: 83.4 tokens/second on Qwen 30B.
- Axelera Server 250p — four AIPUs on a full-height, full-length card. Combined >2,500 TOPS. Published result: up to 4,205 tokens/second on Qwen3 8B, at up to 32.1 tokens/second per watt.
Validated systems on day one: Dell XE5 and Supermicro 111AD. The wider compatibility list includes HPE, Lenovo, Advantech, Axiomtek, 2CRSi and Seco. Dell, integrator E4 Computer Engineering and Axelera are also collaborating on European AI infrastructure projects.
The sales context Axelera publishes: a pipeline exceeding $1.5 billion and more than 600 customers. Next in line is Titania, described as its first chiplet, aimed at rack-mount servers and supercomputers.
The efficiency claims, and why they need care
Axelera publishes several comparison figures, and they are not the same metric, so they should not be merged:
- 3–5× performance-to-cost and 2–3× power efficiency versus leading competitors (Axelera).
- 6× tokens per watt versus GPU approaches, and >8× tokens at equal price versus best competing product (via Chinese tech press reporting of the launch).
- 3–6× versus Nvidia L40 / Jetson (via a separate Chinese analysis piece).
All are vendor-stated and none were independently reproduced in the sources found. The honest reading is: Axelera is arguing that tokens per dollar and tokens per watt — not TOPS — are the right unit for inference economics, which is a defensible position, but the multipliers are marketing figures until someone benchmarks them.
The one figure with a testable shape attached is 32.1 tok/s/W on Qwen3 8B for the Server 250p, because it names a model and a card.
Engineering takeaway
Two things make this worth tracking. First, 629 TOPS inside 45 W on a PCIe card that drops into an existing Dell or Supermicro chassis is a procurement path that does not require a new server architecture — which is often the real barrier for startup silicon. Second, the 16 RISC-V vector cores are a concrete example of the pattern that keeps showing up in AI accelerators: RISC-V as the control and data-movement plane next to a proprietary or custom matrix engine.
If you are evaluating it, benchmark your own model. The published tokens/second figures are single-model, vendor-measured, and the comparison multipliers are not independently verified.
Sources
- Web Pulse — "Axelera Launches Europa AI Processing Unit (AIPU) With Up to 629 TOPS" — https://wpnews.pro/news/axelera-launches-europa-ai-processing-unit-aipu-with-up-to-629-tops
- The Next Gen Tech Insider — "Axelera AI Ships Europa Chip for Dell and Supermicro Server Platforms", citing Axelera AI, TechPowerUp and Datacenter Dynamics — https://www.thenextgentechinsider.com/pulse/axelera-ai-ships-europa-chip-for-dell-and-supermicro-server-platforms
- Remio — "Axelera AI Europa Launch Takes Its Nvidia Alternative Into Dell and Supermicro Servers" — https://www.remio.ai/post/axelera-ai-europa-launch-takes-its-nvidia-alternative-into-dell-and-supermicro-s
- IT之家 (via Tencent News) — "Axelera AI 正式发售 AIPU 芯片 Europa,人工智能算力 629 TOPS", 16 September 2026 — https://news.qq.com/rain/a/20260916A05DCQ00
- 钛媒体 — "Axelera Europa 正式发售:629 TOPS 推理芯片挑战英伟达的 Token 经济学", 15 September 2026 — https://www.tmtpost.com/agent/ai-article?id=20407
Verification notes
- Specs (8 AI cores, 16 RISC-V vector cores, 128 MB L2 SRAM, 256-bit LPDDR5, 200 GB/s, INT4/8/16, 629 TOPS, 45 W TDP, PCIe 4.0 ×4, H.264/H.265 decoder, D-IMC, Voyager SDK, three form factors) are consistent across at least three independent write-ups and are attributed in the body.
- 629 TOPS is quoted at INT8 by TechPowerUp; other sources state 629 TOPS without a precision qualifier. The INT8 attribution is preserved rather than generalised.
- Benchmarks are vendor-reported and single-model: Edge 232p 83.4 tok/s on Qwen 30B; Server 250p 4,205 tok/s on Qwen3 8B at 32.1 tok/s/W. No batch size, quantisation, context length or prompt was published with them. Not independently verified.
- Efficiency multipliers conflict across sources (6× tokens/W vs 2–3× power efficiency vs 3–6× vs L40/Jetson). They are listed separately with attribution and explicitly labelled vendor marketing figures rather than reconciled into one number.
- Samsung 5 nm is described as "according to independent coverage," not an Axelera statement in the sources found.
- Up to 64 GB per chip comes from Remio only; the other sources give bandwidth (200 GB/s) but not capacity.
- $1.5B pipeline and 600+ customers are Axelera press-release claims carried by secondary sources.
- Silicon availability: "now shipping" is supported by Datacenter Dynamics coverage and the validated Dell/Supermicro configurations. Lead times and pricing are not disclosed in any source.
- Overlap note: this site already published a piece on the Europa architecture and Samsung foundry at announcement time (October 2025). This draft covers only the 15 September 2026 shipping milestone and published throughput figures. Editor should decide whether to publish as a follow-up or fold the numbers into the existing article.