Teaching OpenBLAS About the K3's VLEN=1024 AI Cores: HPL on a Heterogeneous RISC-V SoC
The SpacemiT K3 is unusual even by RISC-V standards: one die, two asymmetric RVV 1.0 clusters. Eight X100 general-purpose cores with VLEN=256, and eight A100 "AI" cores with a VLEN=1024 vector unit — the widest vector length shipping in silicon on any RISC-V part. SpacemiT's pitch is that the A100's width exists for the IME2 matrix extension (the vendor's matrix path, rated up to 60 TOPS). But the A100 cores also speak plain RVV, which raises an obvious question the vendor does not answer: can you just throw standard BLAS/HPL workloads at the wide cores and go faster?
The opensolvers team spent time on exactly that question with a Banana Pi BPI-SM10 (K3, Pico-ITX) provided by Banana Pi. Their short answer: no — not with stock VLEN=256 OpenBLAS kernels. Their long answer is more interesting: they added a new RISCV64_ZVL1024B target to OpenBLAS (draft PR #6066), worked around a kernel-side affinity fence, and produced the first public HPL numbers we have seen for both K3 clusters. Once the A100 cores got kernels that match their width, an equal 8+8 mixed configuration reached 57.5 GFLOP/s — beating the X100-only baseline of ~53.
riscv64/generic), FlexiBLAS over OpenBLAS 0.3.34, HPL 2.3, NB=192, OMP_NUM_THREADS=1 unless noted, all residual checks PASSED. The 60 TOPS figure is SpacemiT's vendor rating for the IME2 matrix path — it was not measured in this work, which is explicitly the "plain-RVV story only"; IME2/llama.cpp and GPU numbers are still to come. OpenBLAS PR #6066 is a draft, and its authors deliberately do not propose it as K3's main path.Two Clusters, One Die: X100 vs A100
First, the machine. The BPI-SM10 pairs the K3 SoC (CoM260 module) with an integrated IMG PowerVR BXM-4-64 GPU (Vulkan 1.3 / OpenCL 3.0). The two CPU clusters differ in more than just vector length:
| Property | X100 (cores 0–7) | A100 (cores 8–15) |
|---|---|---|
| Role | General-purpose CPU | AI / matrix (IME2) |
| Clock (typ.) | up to 2.4 GHz | up to 2.0 GHz |
| Clusters | 0–1 (cluster_cpus 0–3, 4–7) | 2–3 (cluster_cpus 8–11, 12–15) |
| ISA (hart) | rv64imafdcvh + bitmanip / vector crypto | rv64imafdcv (no h) + same vector extension family |
| Hypervisor | yes (h) | no |
| RVV | VLEN=256 (vlenb=32, measured) | VLEN=1024 (vlenb=128, measured on cpu8) |
| Custom | — | IME2 · up to 60 TOPS (vendor matrix path) |
| Default Linux scheduling | normal SMP | fenced — online in sysfs, but not in a normal task's affinity mask |
Both clusters implement RVV 1.0 — the A100 is not scalar-only. But per the SM10 notes, plain RVV on the A100 is often slower than on the X100 without vendor IME2 code: the wide VLEN exists for the matrix path, not as a drop-in "faster X60."
Getting Onto the A100 Cores: /proc/set_ai_thread
SpacemiT's kernel keeps general threads on the X100 cluster even when the A100s sit idle. Stock taskset -c 8 fails with EINVAL until the thread is registered as an "AI thread" through a vendor procfs knob:
echo $$ > /proc/set_ai_thread # world-writable (mode 0222) on Bianbu;
# unlocks cores 8-15 for this TID
taskset -c 8-15 ./hpl_binary # affinity now works; the process stays
# on A100 for its lifetime
After the write, the process's Cpus_allowed_list becomes 8–15 — the X100 cluster is no longer available to that task at all. This is a fence, not a load balancer.
And there is a sharper trap lurking underneath, one that will bite anyone porting vector code to this SoC:
exec of the real binary, so the program never runs with a mismatched cached VLEN.For mixed MPI jobs (like HPL), each rank destined for the A100 cluster writes its TID to /proc/set_ai_thread and pins to cores 8–15 before exec.
Kernel Geometry Must Match VLEN
The reason stock OpenBLAS underperforms on the A100 is kernel geometry. Two OpenBLAS targets matter on the K3:
RISCV64_ZVL256Bon the X100 (VLEN=256) — plus a small portability fix, because stock ZVL256B wide-load andvgetLMUL splits assumeVLMAX@256and can corrupt tails when run at a larger VLEN.RISCV64_ZVL1024Bon the A100 (VLEN=1024) — a new target the team generated: DGEMM16×8(plus S/C/Z/TRMM) from OpenBLAS'sgenerate_kernel.py, compiled with-march=..._zvl1024b. Shipped asOpenBLAS-0.3.34_add-riscv64-zvl1024b.patchin the opensolvers/benchmarks repo.
Simply porting ZVL256B kernels to VLEN=1024 without changing the tile shape wastes the machine — and some stock wide-load/vget patterns assume VLMAX@256 outright. One corollary worth engraving: a static TARGET=RISCV64_ZVL1024B build must not run on the X100 cores — use DYNAMIC_ARCH or separate backends per cluster.
HPL Results (N=12000, NB=192, Residuals PASSED)
Single-Cluster Baselines (8 ranks, 2×4 grid)
| Setup | OpenBLAS target | GFLOP/s | Residual |
|---|---|---|---|
| X100 ×8 | ZVL256B | 52.0 | PASSED |
| A100 ×8 | ZVL256B (VLEN-safe) | 15.0 | PASSED |
| A100 ×8 | ZVL1024B (new) | 35.9 | PASSED |
Three numbers tell the whole story. With stock-ish ZVL256B kernels, the A100 cluster runs HPL at 15 GFLOP/s — about 3.5× slower than the X100. Matching the tile geometry to VLEN=1024 recovers roughly 2.4× (35.9 GFLOP/s), which leaves the A100 still about 1.45× behind the X100 in plain RVV mode. For contrast, stock NETLIB BLAS on the A100 managed ~0.8 GFLOP/s at N=4000. At smaller problem sizes, the new target gives 22.5 GFLOP/s (N=4000) and 30.4 GFLOP/s (N=8000) — the wide cores need large N to amortize their geometry.
Equal 16-Rank Mixed (8+8): the Optimum Flips
| N | BLAS on A100 | GFLOP/s | Residual |
|---|---|---|---|
| 4000 | ZVL256B | 16.2 | PASSED |
| 8000 | ZVL256B | 23.7 | PASSED |
| 12000 | ZVL256B | 27.5 | PASSED |
| 12000 | ZVL1024B (X100 still ZVL256B) | 57.5 | PASSED |
This table is the quiet headline. With slow ZVL256B kernels on the A100, giving the AI cores an equal share of ranks is a trap — 8+8 manages only 27.5 GFLOP/s, and the old optimum was "mostly X100, a few A100 ranks." With ZVL1024B on the A100 side, the equal 8+8 configuration becomes the best mixed configuration and beats X100-only outright.
Heterogeneous Sweep (N=12000)
Per-rank FlexiBLAS routes X100 ranks to OPENBLAS_DUAL (ZVL256B) and A100 ranks to OPENBLAS_ZVL1024B, wrapped with HPL_NX100 / HPL_A100_THREADS and the set_ai_thread + taskset dance:
| Config | Ranks | Grid | A100 OMP threads | GFLOP/s | Residual |
|---|---|---|---|---|---|
| 8 X100 + 8 A100 | 8+8 | 4×4 | 1 | 57.5 | PASSED |
| 8 + 4×2 | 8+4 | 3×4 | 2 | 56.2 | PASSED |
| 8 + 2×4 | 8+2 | 2×5 | 4 | 55.5 | PASSED |
| 8 + 1×8 | 8+1 | 3×3 | 8 | 47.6 | PASSED |
| X100 only (ref.) | 8 | 2×4 | — | 53.2 | PASSED |
For reference: an earlier all-ZVL256B heterogeneous run peaked at 51.3 GFLOP/s with 8+2×4. The new A100 target flips the optimum to equal 16 ranks. Full L2/L3 OpenBLAS BLATs (S/D/C/Z) on the A100 with the new target: all PASSED.
Reading the Numbers
Four things these tables establish, independent of any vendor claim:
- Wide VLEN alone is not a free speedup. The A100's VLEN=1024 is slower than VLEN=256 for stock-kernel HPL, because kernel geometry (tile shape, LMUL assumptions,
vgetpatterns) is written against a specific VLEN class. Performance follows the kernels, not the register width. - But matching geometry pays. One new target — a generated DGEMM 16×8 kernel family — recovered 2.4× on the A100 cluster. The remaining 1.45× gap to the X100 is the price of running "plain RVV" on hardware whose width is really meant for IME2.
- The AI cores are useful for ordinary code once the software catches up. 57.5 vs 53.2 GFLOP/s is a ~8% total-system gain from eight cores that default Linux will not even schedule. That is the first public number we have seen demonstrating it in a standard HPC workload.
- Heterogeneous MPI needs per-rank BLAS. The team used FlexiBLAS so X100 and A100 ranks could load different OpenBLAS backends per process — a pattern any K3 HPC/cluster software will need to copy.
Takeaways for Builders
- The K3 is genuinely heterogeneous. X100 remains the primary RVV compute path; the A100's width is there for IME2, not as a "faster X60." Plan software around two clusters, not one 16-core SMP machine — the affinity fence enforces this whether you like it or not.
- Register AI threads before
exec. The/proc/set_ai_thread+tasksetsequence must happen before the dynamic linker caches VLEN decisions; migrating a live process across clusters can SIGSEGV. - Use DYNAMIC_ARCH or per-cluster backends. A static ZVL1024B build will not run on the X100. FlexiBLAS (or equivalent per-process dispatch) is the current answer for mixed jobs.
- PR #6066 is a draft on purpose. The authors explicitly do not propose the ZVL1024B target as K3's main path — treat it as a working patch to pull from the benchmarks repo, and expect the upstream conversation (and the IME2 numbers) to continue.
Related coverage on this site: our annotation of the SpacemiT K3 preview paper (the X100/A100 design intent, vendor-published TOPS tables and scheduling model), the K3 Pico-ITX board itself, MNN LLM inference via IME2 on the K3 (the matrix path this benchmark deliberately avoids), OpenSBI K3 platform support, and the K3's RVA23 architecture.
SpacemiT K3 即使放在 RISC-V 阵营里也相当特殊:一颗芯片、两个非对称的 RVV 1.0 集群——八个 X100 通用核(VLEN=256),加八个 A100「AI 核」,后者带 VLEN=1024 的向量单元,是目前所有 RISC-V 芯片中已量产的最宽向量长度。进迭时空的官方叙事是:A100 的宽度是为 IME2 矩阵扩展(厂商矩阵通路,标称最高 60 TOPS)准备的。但 A100 内核同样支持普通 RVV,这就留下一个厂商没有回答的问题:能不能直接把标准 BLAS/HPL 负载扔给宽核跑出更快的结果?
opensolvers 团队用 Banana Pi 提供的 BPI-SM10(K3,Pico-ITX)专门回答了这个问题。简短答案:不能——原版 VLEN=256 的 OpenBLAS 内核做不到。更长的答案更有价值:他们为 OpenBLAS 新增了 RISCV64_ZVL1024B 目标(草案 PR #6066),绕过了内核侧的亲和性围栏,并给出了我们见过的第一份覆盖 K3 两个集群的公开 HPL 数据。当 A100 拿到与自身宽度匹配的内核后,8+8 均等混部配置达到 57.5 GFLOP/s——反超 X100 单集群约 53 的基线。
riscv64/generic)、FlexiBLAS 套 OpenBLAS 0.3.34、HPL 2.3、NB=192、除注明外 OMP_NUM_THREADS=1,全部残差校验 PASSED。60 TOPS 是进迭时空对 IME2 矩阵通路的厂商标称值——本工作没有测它,作者明确说明这只是「纯 RVV 的故事」;IME2/llama.cpp 与 GPU 数据尚未发布。OpenBLAS PR #6066 是刻意保持的草案,作者明确不提议将其作为 K3 的主线方案。一颗芯片,两个集群:X100 vs A100
先看机器。BPI-SM10 在 K3 SoC(CoM260 模组)之外,还集成了 IMG PowerVR BXM-4-64 GPU(Vulkan 1.3 / OpenCL 3.0)。两个 CPU 集群的差异远不止向量长度:
| 属性 | X100(核 0–7) | A100(核 8–15) |
|---|---|---|
| 角色 | 通用 CPU | AI / 矩阵(IME2) |
| 典型频率 | 最高 2.4 GHz | 最高 2.0 GHz |
| 集群 | 0–1(cluster_cpus 0–3、4–7) | 2–3(cluster_cpus 8–11、12–15) |
| Hart ISA | rv64imafdcvh + bitmanip / 向量加密 | rv64imafdcv(无 h)+ 同族向量扩展 |
| Hypervisor | 支持(h) | 不支持 |
| RVV | VLEN=256(实测 vlenb=32) | VLEN=1024(cpu8 实测 vlenb=128) |
| 自定义扩展 | — | IME2 · 最高 60 TOPS(厂商矩阵通路) |
| 默认 Linux 调度 | 正常 SMP | 围栏隔离——sysfs 可见,但不在普通任务亲和掩码内 |
两个集群都实现了 RVV 1.0——A100 并非纯标量核。但按 SM10 笔记的说法,没有厂商 IME2 代码时,A100 上跑普通 RVV 往往比 X100 更慢:宽 VLEN 是为矩阵通路准备的,不是「更快的 X60」的即插即用替代。
登上 A100 核:/proc/set_ai_thread
进迭时空的内核会把普通线程留在 X100 集群上,哪怕 A100 全部空闲。未经注册的 taskset -c 8 会直接返回 EINVAL——必须先通过厂商的 procfs 接口把线程注册为「AI 线程」:
echo $$ > /proc/set_ai_thread # Bianbu 上全局可写(mode 0222);
# 为本 TID 解锁 8-15 号核
taskset -c 8-15 ./hpl_binary # 亲和性设置生效;进程终身留在 A100
写入之后,进程的 Cpus_allowed_list 变成 8–15——X100 集群对该任务彻底不可见。这是围栏,不是负载均衡。
exec 真正的二进制之前注册 AI 线程,让程序从未以不匹配的缓存 VLEN 运行过。对混合 MPI 作业(如 HPL),每个要落到 A100 的 rank 都在 exec 前把 TID 写入 /proc/set_ai_thread 并绑定到 8–15 号核。
内核几何必须匹配 VLEN
原版 OpenBLAS 在 A100 上跑不动的原因是内核几何。K3 上有两个关键目标:
RISCV64_ZVL256B(X100,VLEN=256)——外加一个小可移植性修复:原版 ZVL256B 的 wide-load 与vgetLMUL 拆分假设VLMAX@256,在更大 VLEN 下运行会损坏尾部数据。RISCV64_ZVL1024B(A100,VLEN=1024)——团队新生成的目标:用 OpenBLAS 的generate_kernel.py生成 DGEMM16×8(含 S/C/Z/TRMM),以-march=..._zvl1024b编译。补丁以OpenBLAS-0.3.34_add-riscv64-zvl1024b.patch的形式发布在 opensolvers/benchmarks 仓库。
不改 tile 形状、直接把 ZVL256B 内核移植到 VLEN=1024 是浪费这台机器——部分原生 wide-load/vget 模式更是直接假设 VLMAX@256。一个值得刻在脑子里的推论:静态 TARGET=RISCV64_ZVL1024B 构建绝不能跑在 X100 上——请用 DYNAMIC_ARCH 或按集群拆分后端。
HPL 结果(N=12000、NB=192、残差全部 PASSED)
单集群基线(8 rank,2×4 网格)
| 配置 | OpenBLAS 目标 | GFLOP/s | 残差 |
|---|---|---|---|
| X100 ×8 | ZVL256B | 52.0 | PASSED |
| A100 ×8 | ZVL256B(VLEN 安全版) | 15.0 | PASSED |
| A100 ×8 | ZVL1024B(新增) | 35.9 | PASSED |
三个数字讲完整个故事。用接近原版的 ZVL256B 内核,A100 集群的 HPL 只有 15 GFLOP/s——比 X100 慢约 3.5 倍。让 tile 几何匹配 VLEN=1024 后找回约 2.4 倍(35.9 GFLOP/s),但在纯 RVV 模式下 A100 仍落后 X100 约 1.45 倍。作为对照:NETLIB 参考实现在 A100 上只有约 0.8 GFLOP/s(N=4000)。更小规模下,新目标给出 22.5 GFLOP/s(N=4000)与 30.4 GFLOP/s(N=8000)——宽核需要足够大的 N 来摊薄几何开销。
均等 16 rank 混部(8+8):最优解翻转
| N | A100 侧 BLAS | GFLOP/s | 残差 |
|---|---|---|---|
| 4000 | ZVL256B | 16.2 | PASSED |
| 8000 | ZVL256B | 23.7 | PASSED |
| 12000 | ZVL256B | 27.5 | PASSED |
| 12000 | ZVL1024B(X100 仍 ZVL256B) | 57.5 | PASSED |
这张表才是安静的头条。A100 用慢速 ZVL256B 时,给 AI 核均等份额是个陷阱——8+8 只有 27.5 GFLOP/s,老的最优解是「X100 为主、A100 少量」。A100 换上 ZVL1024B 后,8+8 均等配置直接成为最优混部方式,并反超 X100 单集群。
异构扫描(N=12000)
FlexiBLAS 按路由分发:X100 rank 用 OPENBLAS_DUAL(ZVL256B),A100 rank 用 OPENBLAS_ZVL1024B,外层包 HPL_NX100 / HPL_A100_THREADS 与 set_ai_thread + taskset:
| 配置 | Rank | 网格 | A100 OMP 线程 | GFLOP/s | 残差 |
|---|---|---|---|---|---|
| 8 X100 + 8 A100 | 8+8 | 4×4 | 1 | 57.5 | PASSED |
| 8 + 4×2 | 8+4 | 3×4 | 2 | 56.2 | PASSED |
| 8 + 2×4 | 8+2 | 2×5 | 4 | 55.5 | PASSED |
| 8 + 1×8 | 8+1 | 3×3 | 8 | 47.6 | PASSED |
| 仅 X100(参照) | 8 | 2×4 | — | 53.2 | PASSED |
参照:更早的全 ZVL256B 异构配置峰值是 8+2×4 时的 51.3 GFLOP/s。新的 A100 目标把最优解翻转为均等 16 rank。新目标下 A100 的全量 L2/L3 OpenBLAS BLATs(S/D/C/Z):全部 PASSED。
怎么读这些数字
- 宽 VLEN 本身不是免费提速。面对原版内核,VLEN=1024 的 A100 反而比 VLEN=256 慢——因为内核几何(tile 形状、LMUL 假设、
vget模式)是针对特定 VLEN 档写的。性能跟着内核走,不跟着寄存器宽度走。 - 但几何匹配就有回报。一个新生成的 DGEMM 16×8 内核家族,就在 A100 集群上找回 2.4 倍。剩下对 X100 的 1.45 倍差距,是在「宽度其实为 IME2 准备」的硬件上跑普通 RVV 的代价。
- 软件跟上来之后,AI 核对普通代码也有用。57.5 vs 53.2 GFLOP/s——默认 Linux 连调度都不肯调度的八个核,给整机带来约 8% 的增益。这是我们见过的第一个在标准 HPC 负载中证明这一点的公开数据。
- 异构 MPI 需要按 rank 配 BLAS。团队用 FlexiBLAS 让 X100 与 A100 的 rank 按进程加载不同 OpenBLAS 后端——任何要做 K3 HPC/集群软件的人都得复刻这个模式。
给开发者的要点
- K3 是真正的异构机器。X100 仍是主 RVV 算力通路;A100 的宽度属于 IME2,不是「更快的 X60」。软件要按两个集群来设计,而不是一台 16 核 SMP——亲和性围栏会替你做这个决定。
- 在
exec之前注册 AI 线程。/proc/set_ai_thread+taskset必须发生在动态链接器缓存 VLEN 决策之前;跨集群迁移活进程可能 SIGSEGV。 - 用 DYNAMIC_ARCH 或按集群拆后端。静态 ZVL1024B 构建跑不了 X100。FlexiBLAS(或等价的按进程分发)是当前混合作业的标准答案。
- PR #6066 是刻意的草案。作者明确不提议将 ZVL1024B 目标作为 K3 主线——把它当作从 benchmarks 仓库拉取的工作补丁即可,上游讨论与 IME2 数据还会继续。
本站相关报道:SpacemiT K3 预印本论文注解(X100/A100 设计意图、厂商 TOPS 表与调度模型)、K3 Pico-ITX 板卡、MNN 在 K3 上经 IME2 做 LLM 推理(本评测刻意避开的矩阵通路)、OpenSBI K3 平台支持补丁、K3 的 RVA23 架构。
Краткое содержание (RU)
Команда opensolvers добавила в OpenBLAS новый целевой профиль RISCV64_ZVL1024B и прогнала HPL на ИИ-ядрах A100 гетерогенной SoC SpacemiT K3 (Banana Pi BPI-SM10): 8 универсальных ядер X100 (VLEN=256, до 52,0 GFLOP/s) плюс 8 ИИ-ядер A100 (VLEN=1024). С «родными» ядрами VLEN=256 кластер A100 выдавал лишь 15 GFLOP/s — в 3,5 раза медленнее X100; новый целевой профиль поднял результат до 35,9 GFLOP/s (в 2,4 раза). При равном смешанном запуске 8+8 (FlexiBLAS, разные бэкенды на кластер) достигнуто 57,5 GFLOP/s — выше базовых 53,2 GFLOP/s только на X100. Ключевые технические детали: ядра A100 «огорожены» от обычного планировщика — нужен echo $$ > /proc/set_ai_thread до taskset -c 8-15; миграция процесса с X100 на A100 после кэширования длины вектора динамическим компоновщиком может завершиться SIGSEGV; статическая сборка ZVL1024B не должна работать на X100 (нужны DYNAMIC_ARCH или раздельные бэкенды). Все цифры — community-бенчмарки (OpenBLAS 0.3.34, HPL 2.3, N=12000, NB=192, остатки PASSED), 60 TOPS — вендорская оценка пути IME2, в этой работе не измерялась. PR OpenBLAS #6066 — намеренно черновой.
Resumen (ES)
El equipo opensolvers añadió un nuevo objetivo RISCV64_ZVL1024B a OpenBLAS y ejecutó HPL en los núcleos de IA A100 del SoC heterogéneo SpacemiT K3 (Banana Pi BPI-SM10): 8 núcleos de propósito general X100 (VLEN=256, hasta 52,0 GFLOP/s) más 8 núcleos A100 (VLEN=1024). Con kernels ZVL256B de serie, el clúster A100 rendía solo 15 GFLOP/s — 3,5 veces más lento que el X100; el nuevo objetivo lo llevó a 35,9 GFLOP/s (×2,4). En una mezcla equilibrada 8+8 (FlexiBLAS, un backend por clúster) se alcanzaron 57,5 GFLOP/s, superando la base de solo-X100 de 53,2. Detalles clave: los núcleos A100 están apartados del planificador normal — hay que registrar el hilo con echo $$ > /proc/set_ai_thread antes de taskset -c 8-15; migrar un proceso de X100 a A100 tras cachear la VLEN puede provocar SIGSEGV; una compilación estática ZVL1024B no debe ejecutarse en X100 (usar DYNAMIC_ARCH o backends separados). Todas las cifras son benchmarks comunitarios (OpenBLAS 0.3.34, HPL 2.3, N=12000, NB=192, residuos PASSED); los 60 TOPS son la cifra del fabricante para la ruta IME2, no medidos aquí. El PR #6066 de OpenBLAS sigue siendo borrador a propósito.
Résumé (FR)
L'équipe opensolvers a ajouté une nouvelle cible RISCV64_ZVL1024B à OpenBLAS et exécuté HPL sur les cœurs IA A100 du SoC hétérogène SpacemiT K3 (Banana Pi BPI-SM10) : 8 cœurs généralistes X100 (VLEN=256, jusqu'à 52,0 GFLOP/s) plus 8 cœurs IA A100 (VLEN=1024). Avec les kernels ZVL256B d'origine, le cluster A100 ne faisait que 15 GFLOP/s — 3,5× plus lent que le X100 ; la nouvelle cible le porte à 35,9 GFLOP/s (×2,4). En mixte équilibré 8+8 (FlexiBLAS, un backend par cluster), la machine atteint 57,5 GFLOP/s, dépassant la référence X100 seul à 53,2. Points techniques : les cœurs A100 sont mis à l'écart de l'ordonnanceur normal — il faut echo $$ > /proc/set_ai_thread avant taskset -c 8-15 ; migrer un processus de X100 vers A100 après la mise en cache de la VLEN peut provoquer un SIGSEGV ; une build statique ZVL1024B ne doit pas tourner sur X100 (DYNAMIC_ARCH ou backends séparés). Tous les chiffres sont des benchmarks communautaires (OpenBLAS 0.3.34, HPL 2.3, N=12000, NB=192, résidus PASSED) ; les 60 TOPS sont la valeur fabricant du chemin IME2, non mesurés ici. La PR OpenBLAS #6066 reste volontairement un brouillon.
Kurzfassung (DE)
Das opensolvers-Team hat OpenBLAS ein neues Ziel RISCV64_ZVL1024B hinzugefügt und HPL auf den KI-Kernen A100 des heterogenen SpacemiT K3 (Banana Pi BPI-SM10) ausgeführt: 8 Allzweckkerne X100 (VLEN=256, bis 52,0 GFLOP/s) plus 8 KI-Kerne A100 (VLEN=1024). Mit Standard-ZVL256B-Kernels lief der A100-Cluster nur mit 15 GFLOP/s — 3,5× langsamer als der X100; das neue Ziel brachte 35,9 GFLOP/s (×2,4). In einer ausgeglichenen 8+8-Mischkonfiguration (FlexiBLAS, ein Backend pro Cluster) wurden 57,5 GFLOP/s erreicht — über der X100-only-Referenz von 53,2. Technisch wichtig: Die A100-Kerne sind vom normalen Scheduler abgeschirmt — zunächst echo $$ > /proc/set_ai_thread, dann taskset -c 8-15; eine Migration von X100 nach A100 nach dem Cachen der VLEN durch den dynamischen Linker kann SIGSEGV auslösen; ein statischer ZVL1024B-Build darf nicht auf X100 laufen (DYNAMIC_ARCH oder getrennte Backends). Alle Zahlen sind Community-Benchmarks (OpenBLAS 0.3.34, HPL 2.3, N=12000, NB=192, Residuen PASSED); die 60 TOPS sind die Herstellerangabe für den IME2-Pfad und wurden hier nicht gemessen. OpenBLAS-PR #6066 bleibt bewusst Entwurf.
خلاصه (FA)
تیم opensolvers هدف جدید RISCV64_ZVL1024B را به OpenBLAS افزود و HPL را روی هستههای هوش مصنوعی A100 تراشه ناهمگن SpacemiT K3 (برد Banana Pi BPI-SM10) اجرا کرد: ۸ هسته عمومی X100 با VLEN=256 (تا 52٫0 گیگافلاپس) بههمراه ۸ هسته A100 با VLEN=1024. با کرنلهای استوک ZVL256B، خوشه A100 فقط ۱۵ گیگافلاپس میگرفت — حدود ۳٫۵ برابر کندتر از X100؛ هدف جدید آن را به ۳۵٫۹ گیگافلاپس رساند (۲٫۴ برابر). در ترکیب متوازن 8+8 (با FlexiBLAS و یک بکاند برای هر خوشه) رقم ۵۷٫۵ گیگافلاپس به دست آمد — بالاتر از خط پایه ۵۳٫۲ گیگافلاپسی فقط-X100. نکات فنی: هستههای A100 از زمانبند عادی جدا هستند — ابتدا echo $$ > /proc/set_ai_thread سپس taskset -c 8-15؛ جابهجایی فرایند از X100 به A100 پس از کششدن VLEN توسط لینکر پویا میتواند SIGSEGV بدهد؛ بیلد استاتیک ZVL1024B نباید روی X100 اجرا شود (DYNAMIC_ARCH یا بکاند جدا). همه اعداد بنچمارکهای جامعه کاربری است (OpenBLAS 0.3.34، HPL 2.3، N=12000، NB=192، باقیماندهها PASSED)؛ عدد ۶۰ TOPS ادعای سازنده برای مسیر IME2 است و اینجا اندازهگیری نشده. PR شماره ۶۰۶۶ OpenBLAS عمداً پیشنویس مانده است.
Sources / 参考来源
- SpacemiT community forum — “SpacemiT K3 has 8× VLEN=1024 AI cores — we taught OpenBLAS about them and ran HPL” by opensolvers (accessed 2026-09-29) — primary post with the HPL tables and takeaways
- opensolvers.com — Banana Pi BPI-SM10 / SpacemiT K3 board notes — full cluster spec table, set_ai_thread mechanism, complete HPL sweep, methodology
- OpenBLAS PR #6066 (draft) — the RISCV64_ZVL1024B target upstreaming attempt
- opensolvers/benchmarks GitHub repo — OpenBLAS-0.3.34_add-riscv64-zvl1024b.patch and the HPL wrappers
- Banana Pi BPI-SM10 (K3-CoM260) product page — the board used (provided by Banana Pi)
- zqb-all/k3_taskset — community tool implementing the set_ai_thread-before-exec pattern
- Our annotation of the SpacemiT K3 preview paper — vendor-published X100/A100 design intent and TOPS tables