Linux RISC-V TLB Flush Path Tracepoint Reaches v2: Debugging Shootdowns on Real Silicon
On 2 October 2026, a v2 revision of [PATCH v2] riscv: mm: Trace TLB flush path selection appeared on the Linux kernel mailing list. Authored by Roman 'Hedin' Storozhenko, the patch adds a RISC-V-specific tracepoint event — riscv_tlb:riscv_tlb_flush_path — so that operators and kernel developers can see which route the kernel chose for every TLB invalidation request: local invalidation, an SBI RFENCE call into M-mode firmware, or a cross-CPU call through Linux's own IPI machinery.
Why TLB Shootdowns Are Hard to Debug on RISC-V
When a process's page tables change, every CPU that may have cached the old mapping needs to invalidate its TLB entries. Linux calls this a TLB shootdown. On RISC-V the kernel can deliver the shootdown through three different mechanisms:
| Path | What happens | When the kernel picks it |
|---|---|---|
local | The current CPU flushes its own TLB only. | The request targets only the local CPU or num_online_cpus() < 2. |
sbi-rfence | Linux delegates the remote fence to the SBI implementation (OpenSBI, vendor firmware, etc.). | riscv_use_sbi_for_rfence() returns true and the hart mask is remote. |
cross-cpu-call | Linux uses its own cross-CPU call / IPI path, typically on_each_cpu(). | Remote invalidation is needed but the SBI RFENCE path is not selected. |
The problem for a kernel engineer is that, until now, the choice was invisible. A slow munmap() or a KVM VM exit could be slow because the SBI firmware is taking a long time, because Linux targeted too many CPUs, or because the request is repeatedly falling back to the slowest path — but ftrace and perf only exposed the generic tlb:tlb_flush event, which records the reason and page count, not the RISC-V-specific path.
The New Tracepoint Event
The patch introduces a new tracepoint class in a new header, include/trace/events/riscv_tlb.h, and instruments flush_tlb_all() and __flush_tlb_range() in arch/riscv/mm/tlbflush.c. The emitted event records:
| Field | Meaning |
|---|---|
start | Virtual address where the invalidation begins. |
size | Size of the invalidated region. |
stride | Invalidation granularity passed to the RISC-V implementation (not forwarded to SBI RFENCE). |
asid | Hardware-visible ASID associated with the request. |
has_mm | Whether a specific mm_struct is attached to the request. |
target_mask_weight | Number of CPUs in the target mask. |
target_cpus | Full target CPU mask, not just the count. |
scope | single, range, address-space, or all. |
path | local, sbi-rfence, or cross-cpu-call. |
The patch deliberately records the full CPU mask because "CPU identity cannot be reconstructed from a count" — if you want to correlate a slow shootdown with scheduler activity, IPI storms, or SBI call latency on a specific hart, you need the mask, not just a weight.
Scope and Path Semantics
The scope is derived from the size/stride and whether an mm is attached:
single:size <= stride.range: a bounded range larger than the stride.address-space:size == FLUSH_TLB_MAX_SIZEandhas_mm == true.all:size == FLUSH_TLB_MAX_SIZEand no mm is attached.
The path is selected at the point the kernel is about to dispatch the invalidation. For flush_tlb_all() the logic is roughly:
if (num_online_cpus() < 2)
trace_riscv_tlb_flush_path(..., path=local);
local_flush_tlb_all();
else if (riscv_use_sbi_for_rfence())
trace_riscv_tlb_flush_path(..., path=sbi-rfence);
sbi_remote_sfence_vma_asid(...);
else
trace_riscv_tlb_flush_path(..., path=cross-cpu-call);
on_each_cpu(__ipi_flush_tlb_all, ...);
How to Use It With ftrace and perf
If the patch lands, a developer on a RISC-V board can enable the event with the standard trace infrastructure:
# enable the event
echo 1 > /sys/kernel/debug/tracing/events/riscv_tlb/riscv_tlb_flush_path/enable
# run the workload that triggers slow TLB shootdowns
...
# read the trace
cat /sys/kernel/debug/tracing/trace | head -n 50
With perf:
perf record -e riscv_tlb:riscv_tlb_flush_path -- sleep 10
perf script
The output format is designed to be parsed:
start=0x7f8c0000 size=0x4000 stride=0x1000 asid=0x12 has_mm=1
target_mask_weight=2 target_cpus=00000000,00000006 scope=range path=sbi-rfence
What This Patch Does Not Solve
- It does not fix slow shootdowns. It only makes the routing observable.
- It does not describe firmware behaviour. When
path=sbi-rfence, the event records that Linux delegated the request to firmware; how OpenSBI or vendor firmware handles it is outside the kernel's visibility. - It is not upstream yet. The reviewer has asked for a concrete use case before accepting the tracepoint.
Builder Takeaways
- TLB shootdown latency is a real portability issue on RISC-V servers. Different boards choose different paths depending on firmware and SBI version; this tracepoint would let you prove which path is taken.
sbi-rfencevscross-cpu-callis the key distinction. If firmware is slow, the SBI path will dominate; if Linux's own path is slow, you have a different tuning target.- Do not script around this until v2 is accepted. The event format and even the existence of the tracepoint may change.
Related on this site: K3 PCIe Root Complex upstream v7, Linux RISC-V patch v8 for Zic64b and RVA23U64 hwprobe, KVM on the K3 reaches stock distros.
2026 年 10 月 2 日,Linux 内核邮件列表出现了 [PATCH v2] riscv: mm: Trace TLB flush path selection 的 v2 版本。补丁作者是 Roman 'Hedin' Storozhenko,它新增了一个 RISC-V 专用的 tracepoint 事件 riscv_tlb:riscv_tlb_flush_path,让运维和内核开发者能看到每一次 TLB 失效请求选择了哪条路径:本地(local)失效、通过 SBI RFENCE 调用 M-mode 固件,还是走 Linux 自己的 跨 CPU 调用(cross-cpu-call) / IPI 机制。
为什么 RISC-V 上的 TLB shootdown 难调试
当进程页表改变时,所有可能缓存了旧映射的 CPU 都需要让对应的 TLB 表项失效,Linux 把这叫 TLB shootdown。在 RISC-V 上,内核可以通过三种机制完成 shootdown:
| 路径 | 行为 | 内核选择条件 |
|---|---|---|
local | 仅当前 CPU 刷新自己的 TLB。 | 请求只针对本地 CPU,或 num_online_cpus() < 2。 |
sbi-rfence | Linux 把远程 fence 委托给 SBI 实现(OpenSBI、厂商固件等)。 | riscv_use_sbi_for_rfence() 返回 true 且目标 mask 是远程的。 |
cross-cpu-call | Linux 使用自己的跨 CPU 调用 / IPI 路径,通常是 on_each_cpu()。 | 需要远程失效但未选择 SBI RFENCE 路径。 |
内核工程师面临的问题在于,此前这个选择是不可见的。一次慢的 munmap() 或 KVM VM exit 之所以慢,可能是因为 SBI 固件耗时太长、Linux 瞄准了过多 CPU,或者请求反复落在最慢路径上——但 ftrace 与 perf 只有通用的 tlb:tlb_flush 事件,它记录 reason 和 page count,不记录 RISC-V 特有的路径。
新的 tracepoint 事件
补丁在新头文件 include/trace/events/riscv_tlb.h 中定义 tracepoint 类,并在 arch/riscv/mm/tlbflush.c 的 flush_tlb_all() 与 __flush_tlb_range() 中埋点。事件记录的字段如下:
| 字段 | 含义 |
|---|---|
start | 失效起始虚拟地址。 |
size | 失效区域大小。 |
stride | 传给 RISC-V 实现的失效粒度(不传给 SBI RFENCE)。 |
asid | 请求关联的硬件可见 ASID。 |
has_mm | 是否有特定的 mm_struct attached。 |
target_mask_weight | 目标 CPU mask 中的 CPU 数量。 |
target_cpus | 完整的目标 CPU mask,不只是数量。 |
scope | single、range、address-space 或 all。 |
path | local、sbi-rfence 或 cross-cpu-call。 |
补丁特意记录完整 CPU mask,因为「仅靠数量无法重建 CPU 身份」——如果你想把一次慢的 shootdown 与调度活动、IPI 风暴或某个 hart 上的 SBI 调用延迟关联起来,就必须拿到 mask。
scope 与 path 语义
scope 由 size/stride 以及是否 attached mm 推导:
single:size <= stride。range:比 stride 大的有限范围。address-space:size == FLUSH_TLB_MAX_SIZE且has_mm == true。all:size == FLUSH_TLB_MAX_SIZE且没有 attached mm。
path 在内核即将分派失效前选定。以 flush_tlb_all() 为例,逻辑大致如下:
if (num_online_cpus() < 2)
trace_riscv_tlb_flush_path(..., path=local);
local_flush_tlb_all();
else if (riscv_use_sbi_for_rfence())
trace_riscv_tlb_flush_path(..., path=sbi-rfence);
sbi_remote_sfence_vma_asid(...);
else
trace_riscv_tlb_flush_path(..., path=cross-cpu-call);
on_each_cpu(__ipi_flush_tlb_all, ...);
配合 ftrace 与 perf 使用
如果补丁合入,RISC-V 板上的开发者可以用标准 trace 基础设施启用该事件:
# 启用事件
echo 1 > /sys/kernel/debug/tracing/events/riscv_tlb/riscv_tlb_flush_path/enable
# 运行触发慢 TLB shootdown 的负载
...
# 读取 trace
cat /sys/kernel/debug/tracing/trace | head -n 50
使用 perf:
perf record -e riscv_tlb:riscv_tlb_flush_path -- sleep 10
perf script
输出格式设计为可解析:
start=0x7f8c0000 size=0x4000 stride=0x1000 asid=0x12 has_mm=1
target_mask_weight=2 target_cpus=00000000,00000006 scope=range path=sbi-rfence
这个补丁不解决什么
- 它不修复慢 shootdown。 只是让路由可观测。
- 它不描述固件行为。 当
path=sbi-rfence时,事件只记录 Linux 把请求委托给了固件;OpenSBI 或厂商固件如何处理不在内核可见范围内。 - 它还没进上游。 评审者要求作者给出具体用例。
工程要点
- TLB shootdown 延迟是 RISC-V 服务器上真实的可移植性问题。 不同板卡根据固件和 SBI 版本选择不同路径;这个 tracepoint 落地后可以用来证明实际走了哪条路。
sbi-rfence与cross-cpu-call是关键区分。 如果固件慢,SBI 路径会占主导;如果 Linux 自己的路径慢,则调优目标不同。- 在 v2 被接受前不要围绕它写脚本。 事件格式乃至 tracepoint 是否存在都可能变化。
本站相关报道:K3 PCIe Root Complex 上游 v7、Linux RISC-V v8 补丁:Zic64b 与 RVA23U64 hwprobe、K3 上的 KVM 进入发行版内核。
Краткое содержание (RU)
2 октября 2026 года Roman 'Hedin' Storozhenko опубликовал v2 патча riscv: mm: Trace TLB flush path selection. Патч добавляет tracepoint riscv_tlb:riscv_tlb_flush_path, который показывает, какой путь выбран для TLB-invalidations: local, SBI RFENCE или cross-CPU call. Событие записывает start, size, stride, ASID, target CPU mask, scope и path. Это v2, всё ещё на рассмотрении; Jisheng Zhang попросил привести реальный пример использования. Пользоваться ftrace/perf, как и любым другим tracepoint.
Resumen (ES)
El 2 de octubre de 2026, Roman 'Hedin' Storozhenko publicó la v2 del parche riscv: mm: Trace TLB flush path selection. Añade el tracepoint riscv_tlb:riscv_tlb_flush_path para observar qué camino elige el kernel para invalidaciones TLB: local, SBI RFENCE o cross-CPU call. El evento registra start, size, stride, ASID, máscara de CPUs destino, scope y path. Es v2 bajo revisión; Jisheng Zhang pidió un caso de uso real. Se puede usar con ftrace/perf.
Résumé (FR)
Le 2 octobre 2026, Roman 'Hedin' Storozhenko a publié la v2 du patch riscv: mm: Trace TLB flush path selection. Il ajoute le tracepoint riscv_tlb:riscv_tlb_flush_path pour observer quel chemin le noyau choisit pour les invalidations TLB : local, SBI RFENCE ou appel inter-CPU. L’événement enregistre start, size, stride, ASID, masque CPU cible, scope et path. Il est en v2 et en revue ; Jisheng Zhang a demandé un cas d’usage concret. Utilisable avec ftrace/perf.
Kurzfassung (DE)
Am 2. Oktober 2026 veröffentlichte Roman 'Hedin' Storozhenko v2 des Patches riscv: mm: Trace TLB flush path selection. Er fügt den Tracepoint riscv_tlb:riscv_tlb_flush_path hinzu, um zu beobachten, welchen Weg der Kernel für TLB-Invalidierungen wählt: local, SBI RFENCE oder Cross-CPU-Call. Das Ereignis zeichnet start, size, stride, ASID, Ziel-CPU-Mask, scope und path auf. Es ist v2 und unter Review; Jisheng Zhang bat um ein konkretes Anwendungsbeispiel. Verwendbar mit ftrace/perf.
خلاصه (FA)
در ۲ اکتبر ۲۰۲۶، Roman 'Hedin' Storozhenko نسخهی v2 پچ riscv: mm: Trace TLB flush path selection را منتشر کرد. این پچ tracepoint riscv_tlb:riscv_tlb_flush_path را اضافه میکند تا مشخص شود کرده TLB کدام مسیر را انتخاب کرده: local، SBI RFENCE یا cross-CPU call. رخداد start، size، stride، ASID، target CPU mask، scope و path را شامل میکند. هنوز v2 است و در حال بررسی، Jisheng Zhang از متوع یک مورد کاربردی کنکرت خواست. قابل استفاده با ftrace/perf.
Sources / 参考来源
- LKML archive — Jisheng Zhang, "Re: [PATCH v2] riscv: mm: Trace TLB flush path selection", Fri Oct 02 2026 12:04:47 EST (review reply quoting Roman Storozhenko's v2 patch)
- Lore.kernel.org — v1 of the same patch series (Aug 29 2026)
- RISC-V specifications — supervisor counter-enable and SBI RFENCE semantics
- OpenSBI repository — reference SBI implementation that handles RFENCE calls