Software · HPC · Benchmarks

Teaching OpenBLAS About the K3's VLEN=1024 AI Cores: HPL on a Heterogeneous RISC-V SoC

📅 2026-09-29 ⏱ 11 min read 🔗 Source: SpacemiT community forum (opensolvers), opensolvers.com board notes, OpenBLAS PR #6066

The SpacemiT K3 is unusual even by RISC-V standards: one die, two asymmetric RVV 1.0 clusters. Eight X100 general-purpose cores with VLEN=256, and eight A100 "AI" cores with a VLEN=1024 vector unit — the widest vector length shipping in silicon on any RISC-V part. SpacemiT's pitch is that the A100's width exists for the IME2 matrix extension (the vendor's matrix path, rated up to 60 TOPS). But the A100 cores also speak plain RVV, which raises an obvious question the vendor does not answer: can you just throw standard BLAS/HPL workloads at the wide cores and go faster?

The opensolvers team spent time on exactly that question with a Banana Pi BPI-SM10 (K3, Pico-ITX) provided by Banana Pi. Their short answer: no — not with stock VLEN=256 OpenBLAS kernels. Their long answer is more interesting: they added a new RISCV64_ZVL1024B target to OpenBLAS (draft PR #6066), worked around a kernel-side affinity fence, and produced the first public HPL numbers we have seen for both K3 clusters. Once the A100 cores got kernels that match their width, an equal 8+8 mixed configuration reached 57.5 GFLOP/s — beating the X100-only baseline of ~53.

Honesty note. All numbers below are community-published benchmarks from the opensolvers team, not vendor figures, and we have not reproduced them on our own hardware. The board was provided by Banana Pi. The forum post and board write-up do not carry an absolute publication date; we accessed both on 2026-09-29. Software context: Bianbu 4.0.6, EESSI 2025.06-001 (riscv64/generic), FlexiBLAS over OpenBLAS 0.3.34, HPL 2.3, NB=192, OMP_NUM_THREADS=1 unless noted, all residual checks PASSED. The 60 TOPS figure is SpacemiT's vendor rating for the IME2 matrix path — it was not measured in this work, which is explicitly the "plain-RVV story only"; IME2/llama.cpp and GPU numbers are still to come. OpenBLAS PR #6066 is a draft, and its authors deliberately do not propose it as K3's main path.

Two Clusters, One Die: X100 vs A100

First, the machine. The BPI-SM10 pairs the K3 SoC (CoM260 module) with an integrated IMG PowerVR BXM-4-64 GPU (Vulkan 1.3 / OpenCL 3.0). The two CPU clusters differ in more than just vector length:

PropertyX100 (cores 0–7)A100 (cores 8–15)
RoleGeneral-purpose CPUAI / matrix (IME2)
Clock (typ.)up to 2.4 GHzup to 2.0 GHz
Clusters0–1 (cluster_cpus 0–3, 4–7)2–3 (cluster_cpus 8–11, 12–15)
ISA (hart)rv64imafdcvh + bitmanip / vector cryptorv64imafdcv (no h) + same vector extension family
Hypervisoryes (h)no
RVVVLEN=256 (vlenb=32, measured)VLEN=1024 (vlenb=128, measured on cpu8)
Custom—IME2 · up to 60 TOPS (vendor matrix path)
Default Linux schedulingnormal SMPfenced — online in sysfs, but not in a normal task's affinity mask

Both clusters implement RVV 1.0 — the A100 is not scalar-only. But per the SM10 notes, plain RVV on the A100 is often slower than on the X100 without vendor IME2 code: the wide VLEN exists for the matrix path, not as a drop-in "faster X60."

Getting Onto the A100 Cores: /proc/set_ai_thread

SpacemiT's kernel keeps general threads on the X100 cluster even when the A100s sit idle. Stock taskset -c 8 fails with EINVAL until the thread is registered as an "AI thread" through a vendor procfs knob:

echo $$ > /proc/set_ai_thread   # world-writable (mode 0222) on Bianbu;
                                # unlocks cores 8-15 for this TID
taskset -c 8-15 ./hpl_binary    # affinity now works; the process stays
                                # on A100 for its lifetime

After the write, the process's Cpus_allowed_list becomes 8–15 — the X100 cluster is no longer available to that task at all. This is a fence, not a load balancer.

And there is a sharper trap lurking underneath, one that will bite anyone porting vector code to this SoC:

The VLEN-migration pitfall. Do not start a process on the X100 cores (VLEN=256), let the dynamic linker / OpenBLAS cache their vector-length decisions, and then migrate it to the A100 (VLEN=1024). That pattern can SIGSEGV. The opensolvers team's rule — the same pattern used by SpacemiT's AI runtimes and the community k3_taskset tool — is to register the AI thread before exec of the real binary, so the program never runs with a mismatched cached VLEN.

For mixed MPI jobs (like HPL), each rank destined for the A100 cluster writes its TID to /proc/set_ai_thread and pins to cores 8–15 before exec.

Kernel Geometry Must Match VLEN

The reason stock OpenBLAS underperforms on the A100 is kernel geometry. Two OpenBLAS targets matter on the K3:

Simply porting ZVL256B kernels to VLEN=1024 without changing the tile shape wastes the machine — and some stock wide-load/vget patterns assume VLMAX@256 outright. One corollary worth engraving: a static TARGET=RISCV64_ZVL1024B build must not run on the X100 cores — use DYNAMIC_ARCH or separate backends per cluster.

HPL Results (N=12000, NB=192, Residuals PASSED)

Single-Cluster Baselines (8 ranks, 2×4 grid)

SetupOpenBLAS targetGFLOP/sResidual
X100 ×8ZVL256B52.0PASSED
A100 ×8ZVL256B (VLEN-safe)15.0PASSED
A100 ×8ZVL1024B (new)35.9PASSED

Three numbers tell the whole story. With stock-ish ZVL256B kernels, the A100 cluster runs HPL at 15 GFLOP/s — about 3.5× slower than the X100. Matching the tile geometry to VLEN=1024 recovers roughly 2.4× (35.9 GFLOP/s), which leaves the A100 still about 1.45× behind the X100 in plain RVV mode. For contrast, stock NETLIB BLAS on the A100 managed ~0.8 GFLOP/s at N=4000. At smaller problem sizes, the new target gives 22.5 GFLOP/s (N=4000) and 30.4 GFLOP/s (N=8000) — the wide cores need large N to amortize their geometry.

Equal 16-Rank Mixed (8+8): the Optimum Flips

NBLAS on A100GFLOP/sResidual
4000ZVL256B16.2PASSED
8000ZVL256B23.7PASSED
12000ZVL256B27.5PASSED
12000ZVL1024B (X100 still ZVL256B)57.5PASSED

This table is the quiet headline. With slow ZVL256B kernels on the A100, giving the AI cores an equal share of ranks is a trap — 8+8 manages only 27.5 GFLOP/s, and the old optimum was "mostly X100, a few A100 ranks." With ZVL1024B on the A100 side, the equal 8+8 configuration becomes the best mixed configuration and beats X100-only outright.

Heterogeneous Sweep (N=12000)

Per-rank FlexiBLAS routes X100 ranks to OPENBLAS_DUAL (ZVL256B) and A100 ranks to OPENBLAS_ZVL1024B, wrapped with HPL_NX100 / HPL_A100_THREADS and the set_ai_thread + taskset dance:

ConfigRanksGridA100 OMP threadsGFLOP/sResidual
8 X100 + 8 A1008+84×4157.5PASSED
8 + 4×28+43×4256.2PASSED
8 + 2×48+22×5455.5PASSED
8 + 1×88+13×3847.6PASSED
X100 only (ref.)82×4—53.2PASSED

For reference: an earlier all-ZVL256B heterogeneous run peaked at 51.3 GFLOP/s with 8+2×4. The new A100 target flips the optimum to equal 16 ranks. Full L2/L3 OpenBLAS BLATs (S/D/C/Z) on the A100 with the new target: all PASSED.

Reading the Numbers

Four things these tables establish, independent of any vendor claim:

Takeaways for Builders

  1. The K3 is genuinely heterogeneous. X100 remains the primary RVV compute path; the A100's width is there for IME2, not as a "faster X60." Plan software around two clusters, not one 16-core SMP machine — the affinity fence enforces this whether you like it or not.
  2. Register AI threads before exec. The /proc/set_ai_thread + taskset sequence must happen before the dynamic linker caches VLEN decisions; migrating a live process across clusters can SIGSEGV.
  3. Use DYNAMIC_ARCH or per-cluster backends. A static ZVL1024B build will not run on the X100. FlexiBLAS (or equivalent per-process dispatch) is the current answer for mixed jobs.
  4. PR #6066 is a draft on purpose. The authors explicitly do not propose the ZVL1024B target as K3's main path — treat it as a working patch to pull from the benchmarks repo, and expect the upstream conversation (and the IME2 numbers) to continue.

Related coverage on this site: our annotation of the SpacemiT K3 preview paper (the X100/A100 design intent, vendor-published TOPS tables and scheduling model), the K3 Pico-ITX board itself, MNN LLM inference via IME2 on the K3 (the matrix path this benchmark deliberately avoids), OpenSBI K3 platform support, and the K3's RVA23 architecture.

SpacemiT K3 即使放在 RISC-V 阵营里也相当特殊:一颗芯片、两个非对称的 RVV 1.0 集群——八个 X100 通用核(VLEN=256),加八个 A100「AI 核」,后者带 VLEN=1024 的向量单元,是目前所有 RISC-V 芯片中已量产的最宽向量长度。进迭时空的官方叙事是:A100 的宽度是为 IME2 矩阵扩展(厂商矩阵通路,标称最高 60 TOPS)准备的。但 A100 内核同样支持普通 RVV,这就留下一个厂商没有回答的问题:能不能直接把标准 BLAS/HPL 负载扔给宽核跑出更快的结果?

opensolvers 团队用 Banana Pi 提供的 BPI-SM10(K3,Pico-ITX)专门回答了这个问题。简短答案:不能——原版 VLEN=256 的 OpenBLAS 内核做不到。更长的答案更有价值:他们为 OpenBLAS 新增了 RISCV64_ZVL1024B 目标(草案 PR #6066),绕过了内核侧的亲和性围栏,并给出了我们见过的第一份覆盖 K3 两个集群的公开 HPL 数据。当 A100 拿到与自身宽度匹配的内核后,8+8 均等混部配置达到 57.5 GFLOP/s——反超 X100 单集群约 53 的基线。

诚实声明。 以下全部数字来自 opensolvers 团队的社区公开实测,不是厂商口径,我们也没有在自己的硬件上复现。板卡由 Banana Pi 提供。论坛帖与板卡笔记页均未标注绝对发布日期,本站于 2026-09-29 访问确认内容在册(当时 18 次浏览)。软件环境:Bianbu 4.0.6、EESSI 2025.06-001(riscv64/generic)、FlexiBLAS 套 OpenBLAS 0.3.34、HPL 2.3、NB=192、除注明外 OMP_NUM_THREADS=1,全部残差校验 PASSED。60 TOPS 是进迭时空对 IME2 矩阵通路的厂商标称值——本工作没有测它,作者明确说明这只是「纯 RVV 的故事」;IME2/llama.cpp 与 GPU 数据尚未发布。OpenBLAS PR #6066 是刻意保持的草案,作者明确不提议将其作为 K3 的主线方案。

一颗芯片,两个集群:X100 vs A100

先看机器。BPI-SM10 在 K3 SoC(CoM260 模组)之外,还集成了 IMG PowerVR BXM-4-64 GPU(Vulkan 1.3 / OpenCL 3.0)。两个 CPU 集群的差异远不止向量长度:

属性X100(核 0–7)A100(核 8–15)
角色通用 CPUAI / 矩阵(IME2)
典型频率最高 2.4 GHz最高 2.0 GHz
集群0–1(cluster_cpus 0–3、4–7)2–3(cluster_cpus 8–11、12–15)
Hart ISArv64imafdcvh + bitmanip / 向量加密rv64imafdcv(无 h)+ 同族向量扩展
Hypervisor支持(h)不支持
RVVVLEN=256(实测 vlenb=32)VLEN=1024(cpu8 实测 vlenb=128)
自定义扩展—IME2 · 最高 60 TOPS(厂商矩阵通路)
默认 Linux 调度正常 SMP围栏隔离——sysfs 可见,但不在普通任务亲和掩码内

两个集群都实现了 RVV 1.0——A100 并非纯标量核。但按 SM10 笔记的说法,没有厂商 IME2 代码时,A100 上跑普通 RVV 往往比 X100 更慢:宽 VLEN 是为矩阵通路准备的,不是「更快的 X60」的即插即用替代。

登上 A100 核:/proc/set_ai_thread

进迭时空的内核会把普通线程留在 X100 集群上,哪怕 A100 全部空闲。未经注册的 taskset -c 8 会直接返回 EINVAL——必须先通过厂商的 procfs 接口把线程注册为「AI 线程」:

echo $$ > /proc/set_ai_thread   # Bianbu 上全局可写(mode 0222);
                                # 为本 TID 解锁 8-15 号核
taskset -c 8-15 ./hpl_binary    # 亲和性设置生效;进程终身留在 A100

写入之后,进程的 Cpus_allowed_list 变成 8–15——X100 集群对该任务彻底不可见。这是围栏,不是负载均衡。

VLEN 迁移陷阱。千万不要让进程先在 X100(VLEN=256)上启动、让动态链接器 / OpenBLAS 缓存向量长度决策、然后再迁移到 A100(VLEN=1024)——这个模式可能 SIGSEGV。opensolvers 的规则与 SpacemiT AI 运行时及社区 k3_taskset 工具一致:在 exec 真正的二进制之前注册 AI 线程,让程序从未以不匹配的缓存 VLEN 运行过。

对混合 MPI 作业(如 HPL),每个要落到 A100 的 rank 都在 exec 前把 TID 写入 /proc/set_ai_thread 并绑定到 8–15 号核。

内核几何必须匹配 VLEN

原版 OpenBLAS 在 A100 上跑不动的原因是内核几何。K3 上有两个关键目标:

不改 tile 形状、直接把 ZVL256B 内核移植到 VLEN=1024 是浪费这台机器——部分原生 wide-load/vget 模式更是直接假设 VLMAX@256。一个值得刻在脑子里的推论:静态 TARGET=RISCV64_ZVL1024B 构建绝不能跑在 X100 上——请用 DYNAMIC_ARCH 或按集群拆分后端。

HPL 结果(N=12000、NB=192、残差全部 PASSED)

单集群基线(8 rank,2×4 网格)

配置OpenBLAS 目标GFLOP/s残差
X100 ×8ZVL256B52.0PASSED
A100 ×8ZVL256B(VLEN 安全版)15.0PASSED
A100 ×8ZVL1024B(新增)35.9PASSED

三个数字讲完整个故事。用接近原版的 ZVL256B 内核,A100 集群的 HPL 只有 15 GFLOP/s——比 X100 慢约 3.5 倍。让 tile 几何匹配 VLEN=1024 后找回约 2.4 倍(35.9 GFLOP/s),但在纯 RVV 模式下 A100 仍落后 X100 约 1.45 倍。作为对照:NETLIB 参考实现在 A100 上只有约 0.8 GFLOP/s(N=4000)。更小规模下,新目标给出 22.5 GFLOP/s(N=4000)与 30.4 GFLOP/s(N=8000)——宽核需要足够大的 N 来摊薄几何开销。

均等 16 rank 混部(8+8):最优解翻转

NA100 侧 BLASGFLOP/s残差
4000ZVL256B16.2PASSED
8000ZVL256B23.7PASSED
12000ZVL256B27.5PASSED
12000ZVL1024B(X100 仍 ZVL256B)57.5PASSED

这张表才是安静的头条。A100 用慢速 ZVL256B 时,给 AI 核均等份额是个陷阱——8+8 只有 27.5 GFLOP/s,老的最优解是「X100 为主、A100 少量」。A100 换上 ZVL1024B 后,8+8 均等配置直接成为最优混部方式,并反超 X100 单集群。

异构扫描(N=12000)

FlexiBLAS 按路由分发:X100 rank 用 OPENBLAS_DUAL(ZVL256B),A100 rank 用 OPENBLAS_ZVL1024B,外层包 HPL_NX100 / HPL_A100_THREADS 与 set_ai_thread + taskset:

配置Rank网格A100 OMP 线程GFLOP/s残差
8 X100 + 8 A1008+84×4157.5PASSED
8 + 4×28+43×4256.2PASSED
8 + 2×48+22×5455.5PASSED
8 + 1×88+13×3847.6PASSED
仅 X100(参照)82×4—53.2PASSED

参照:更早的全 ZVL256B 异构配置峰值是 8+2×4 时的 51.3 GFLOP/s。新的 A100 目标把最优解翻转为均等 16 rank。新目标下 A100 的全量 L2/L3 OpenBLAS BLATs(S/D/C/Z):全部 PASSED。

怎么读这些数字

给开发者的要点

  1. K3 是真正的异构机器。X100 仍是主 RVV 算力通路;A100 的宽度属于 IME2,不是「更快的 X60」。软件要按两个集群来设计,而不是一台 16 核 SMP——亲和性围栏会替你做这个决定。
  2. 在 exec 之前注册 AI 线程。/proc/set_ai_thread + taskset 必须发生在动态链接器缓存 VLEN 决策之前;跨集群迁移活进程可能 SIGSEGV。
  3. 用 DYNAMIC_ARCH 或按集群拆后端。静态 ZVL1024B 构建跑不了 X100。FlexiBLAS(或等价的按进程分发)是当前混合作业的标准答案。
  4. PR #6066 是刻意的草案。作者明确不提议将 ZVL1024B 目标作为 K3 主线——把它当作从 benchmarks 仓库拉取的工作补丁即可,上游讨论与 IME2 数据还会继续。

本站相关报道:SpacemiT K3 预印本论文注解(X100/A100 设计意图、厂商 TOPS 表与调度模型)、K3 Pico-ITX 板卡、MNN 在 K3 上经 IME2 做 LLM 推理(本评测刻意避开的矩阵通路)、OpenSBI K3 平台支持补丁、K3 的 RVA23 架构。

Краткое содержание (RU)

Команда opensolvers добавила в OpenBLAS новый целевой профиль RISCV64_ZVL1024B и прогнала HPL на ИИ-ядрах A100 гетерогенной SoC SpacemiT K3 (Banana Pi BPI-SM10): 8 универсальных ядер X100 (VLEN=256, до 52,0 GFLOP/s) плюс 8 ИИ-ядер A100 (VLEN=1024). С «родными» ядрами VLEN=256 кластер A100 выдавал лишь 15 GFLOP/s — в 3,5 раза медленнее X100; новый целевой профиль поднял результат до 35,9 GFLOP/s (в 2,4 раза). При равном смешанном запуске 8+8 (FlexiBLAS, разные бэкенды на кластер) достигнуто 57,5 GFLOP/s — выше базовых 53,2 GFLOP/s только на X100. Ключевые технические детали: ядра A100 «огорожены» от обычного планировщика — нужен echo $$ > /proc/set_ai_thread до taskset -c 8-15; миграция процесса с X100 на A100 после кэширования длины вектора динамическим компоновщиком может завершиться SIGSEGV; статическая сборка ZVL1024B не должна работать на X100 (нужны DYNAMIC_ARCH или раздельные бэкенды). Все цифры — community-бенчмарки (OpenBLAS 0.3.34, HPL 2.3, N=12000, NB=192, остатки PASSED), 60 TOPS — вендорская оценка пути IME2, в этой работе не измерялась. PR OpenBLAS #6066 — намеренно черновой.

Resumen (ES)

El equipo opensolvers añadió un nuevo objetivo RISCV64_ZVL1024B a OpenBLAS y ejecutó HPL en los núcleos de IA A100 del SoC heterogéneo SpacemiT K3 (Banana Pi BPI-SM10): 8 núcleos de propósito general X100 (VLEN=256, hasta 52,0 GFLOP/s) más 8 núcleos A100 (VLEN=1024). Con kernels ZVL256B de serie, el clúster A100 rendía solo 15 GFLOP/s — 3,5 veces más lento que el X100; el nuevo objetivo lo llevó a 35,9 GFLOP/s (×2,4). En una mezcla equilibrada 8+8 (FlexiBLAS, un backend por clúster) se alcanzaron 57,5 GFLOP/s, superando la base de solo-X100 de 53,2. Detalles clave: los núcleos A100 están apartados del planificador normal — hay que registrar el hilo con echo $$ > /proc/set_ai_thread antes de taskset -c 8-15; migrar un proceso de X100 a A100 tras cachear la VLEN puede provocar SIGSEGV; una compilación estática ZVL1024B no debe ejecutarse en X100 (usar DYNAMIC_ARCH o backends separados). Todas las cifras son benchmarks comunitarios (OpenBLAS 0.3.34, HPL 2.3, N=12000, NB=192, residuos PASSED); los 60 TOPS son la cifra del fabricante para la ruta IME2, no medidos aquí. El PR #6066 de OpenBLAS sigue siendo borrador a propósito.

Résumé (FR)

L'équipe opensolvers a ajouté une nouvelle cible RISCV64_ZVL1024B à OpenBLAS et exécuté HPL sur les cœurs IA A100 du SoC hétérogène SpacemiT K3 (Banana Pi BPI-SM10) : 8 cœurs généralistes X100 (VLEN=256, jusqu'à 52,0 GFLOP/s) plus 8 cœurs IA A100 (VLEN=1024). Avec les kernels ZVL256B d'origine, le cluster A100 ne faisait que 15 GFLOP/s — 3,5× plus lent que le X100 ; la nouvelle cible le porte à 35,9 GFLOP/s (×2,4). En mixte équilibré 8+8 (FlexiBLAS, un backend par cluster), la machine atteint 57,5 GFLOP/s, dépassant la référence X100 seul à 53,2. Points techniques : les cœurs A100 sont mis à l'écart de l'ordonnanceur normal — il faut echo $$ > /proc/set_ai_thread avant taskset -c 8-15 ; migrer un processus de X100 vers A100 après la mise en cache de la VLEN peut provoquer un SIGSEGV ; une build statique ZVL1024B ne doit pas tourner sur X100 (DYNAMIC_ARCH ou backends séparés). Tous les chiffres sont des benchmarks communautaires (OpenBLAS 0.3.34, HPL 2.3, N=12000, NB=192, résidus PASSED) ; les 60 TOPS sont la valeur fabricant du chemin IME2, non mesurés ici. La PR OpenBLAS #6066 reste volontairement un brouillon.

Kurzfassung (DE)

Das opensolvers-Team hat OpenBLAS ein neues Ziel RISCV64_ZVL1024B hinzugefügt und HPL auf den KI-Kernen A100 des heterogenen SpacemiT K3 (Banana Pi BPI-SM10) ausgeführt: 8 Allzweckkerne X100 (VLEN=256, bis 52,0 GFLOP/s) plus 8 KI-Kerne A100 (VLEN=1024). Mit Standard-ZVL256B-Kernels lief der A100-Cluster nur mit 15 GFLOP/s — 3,5× langsamer als der X100; das neue Ziel brachte 35,9 GFLOP/s (×2,4). In einer ausgeglichenen 8+8-Mischkonfiguration (FlexiBLAS, ein Backend pro Cluster) wurden 57,5 GFLOP/s erreicht — über der X100-only-Referenz von 53,2. Technisch wichtig: Die A100-Kerne sind vom normalen Scheduler abgeschirmt — zunächst echo $$ > /proc/set_ai_thread, dann taskset -c 8-15; eine Migration von X100 nach A100 nach dem Cachen der VLEN durch den dynamischen Linker kann SIGSEGV auslösen; ein statischer ZVL1024B-Build darf nicht auf X100 laufen (DYNAMIC_ARCH oder getrennte Backends). Alle Zahlen sind Community-Benchmarks (OpenBLAS 0.3.34, HPL 2.3, N=12000, NB=192, Residuen PASSED); die 60 TOPS sind die Herstellerangabe für den IME2-Pfad und wurden hier nicht gemessen. OpenBLAS-PR #6066 bleibt bewusst Entwurf.

خلاصه (FA)

تیم opensolvers هدف جدید RISCV64_ZVL1024B را به OpenBLAS افزود و HPL را روی هسته‌های هوش مصنوعی A100 تراشه ناهمگن SpacemiT K3 (برد Banana Pi BPI-SM10) اجرا کرد: ۸ هسته عمومی X100 با VLEN=256 (تا 52٫0 گیگافلاپس) به‌همراه ۸ هسته A100 با VLEN=1024. با کرنل‌های استوک ZVL256B، خوشه A100 فقط ۱۵ گیگافلاپس می‌گرفت — حدود ۳٫۵ برابر کندتر از X100؛ هدف جدید آن را به ۳۵٫۹ گیگافلاپس رساند (۲٫۴ برابر). در ترکیب متوازن 8+8 (با FlexiBLAS و یک بک‌اند برای هر خوشه) رقم ۵۷٫۵ گیگافلاپس به دست آمد — بالاتر از خط پایه ۵۳٫۲ گیگافلاپسی فقط-X100. نکات فنی: هسته‌های A100 از زمان‌بند عادی جدا هستند — ابتدا echo $$ > /proc/set_ai_thread سپس taskset -c 8-15؛ جابه‌جایی فرایند از X100 به A100 پس از کش‌شدن VLEN توسط لینکر پویا می‌تواند SIGSEGV بدهد؛ بیلد استاتیک ZVL1024B نباید روی X100 اجرا شود (DYNAMIC_ARCH یا بک‌اند جدا). همه اعداد بنچمارک‌های جامعه کاربری است (OpenBLAS 0.3.34، HPL 2.3، N=12000، NB=192، باقیمانده‌ها PASSED)؛ عدد ۶۰ TOPS ادعای سازنده برای مسیر IME2 است و اینجا اندازه‌گیری نشده. PR شماره ۶۰۶۶ OpenBLAS عمداً پیش‌نویس مانده است.

Sources / 参考来源

← Back to Tech Blog