RISC-V Development

Linux RISC-V TLB Flush Path Tracepoint Reaches v2: Debugging Shootdowns on Real Silicon

📅 2026-10-03 ⏱ 7 min read 🔗 Source: LKML / linux-riscv (Oct 2 2026)

On 2 October 2026, a v2 revision of [PATCH v2] riscv: mm: Trace TLB flush path selection appeared on the Linux kernel mailing list. Authored by Roman 'Hedin' Storozhenko, the patch adds a RISC-V-specific tracepoint event — riscv_tlb:riscv_tlb_flush_path — so that operators and kernel developers can see which route the kernel chose for every TLB invalidation request: local invalidation, an SBI RFENCE call into M-mode firmware, or a cross-CPU call through Linux's own IPI machinery.

Honesty note. This is a v2 patch under review, not a merged kernel feature. Jisheng Zhang replied on October 2 asking the author to show a concrete bug or performance problem that only this tracepoint can solve, and suggesting kprobes as an alternative. Treat the event format and field list as provisional.

Why TLB Shootdowns Are Hard to Debug on RISC-V

When a process's page tables change, every CPU that may have cached the old mapping needs to invalidate its TLB entries. Linux calls this a TLB shootdown. On RISC-V the kernel can deliver the shootdown through three different mechanisms:

PathWhat happensWhen the kernel picks it
localThe current CPU flushes its own TLB only.The request targets only the local CPU or num_online_cpus() < 2.
sbi-rfenceLinux delegates the remote fence to the SBI implementation (OpenSBI, vendor firmware, etc.).riscv_use_sbi_for_rfence() returns true and the hart mask is remote.
cross-cpu-callLinux uses its own cross-CPU call / IPI path, typically on_each_cpu().Remote invalidation is needed but the SBI RFENCE path is not selected.

The problem for a kernel engineer is that, until now, the choice was invisible. A slow munmap() or a KVM VM exit could be slow because the SBI firmware is taking a long time, because Linux targeted too many CPUs, or because the request is repeatedly falling back to the slowest path — but ftrace and perf only exposed the generic tlb:tlb_flush event, which records the reason and page count, not the RISC-V-specific path.

The New Tracepoint Event

The patch introduces a new tracepoint class in a new header, include/trace/events/riscv_tlb.h, and instruments flush_tlb_all() and __flush_tlb_range() in arch/riscv/mm/tlbflush.c. The emitted event records:

FieldMeaning
startVirtual address where the invalidation begins.
sizeSize of the invalidated region.
strideInvalidation granularity passed to the RISC-V implementation (not forwarded to SBI RFENCE).
asidHardware-visible ASID associated with the request.
has_mmWhether a specific mm_struct is attached to the request.
target_mask_weightNumber of CPUs in the target mask.
target_cpusFull target CPU mask, not just the count.
scopesingle, range, address-space, or all.
pathlocal, sbi-rfence, or cross-cpu-call.

The patch deliberately records the full CPU mask because "CPU identity cannot be reconstructed from a count" — if you want to correlate a slow shootdown with scheduler activity, IPI storms, or SBI call latency on a specific hart, you need the mask, not just a weight.

Scope and Path Semantics

The scope is derived from the size/stride and whether an mm is attached:

The path is selected at the point the kernel is about to dispatch the invalidation. For flush_tlb_all() the logic is roughly:

if (num_online_cpus() < 2)
    trace_riscv_tlb_flush_path(..., path=local);
    local_flush_tlb_all();
else if (riscv_use_sbi_for_rfence())
    trace_riscv_tlb_flush_path(..., path=sbi-rfence);
    sbi_remote_sfence_vma_asid(...);
else
    trace_riscv_tlb_flush_path(..., path=cross-cpu-call);
    on_each_cpu(__ipi_flush_tlb_all, ...);

How to Use It With ftrace and perf

If the patch lands, a developer on a RISC-V board can enable the event with the standard trace infrastructure:

# enable the event
echo 1 > /sys/kernel/debug/tracing/events/riscv_tlb/riscv_tlb_flush_path/enable

# run the workload that triggers slow TLB shootdowns
...

# read the trace
cat /sys/kernel/debug/tracing/trace | head -n 50

With perf:

perf record -e riscv_tlb:riscv_tlb_flush_path -- sleep 10
perf script

The output format is designed to be parsed:

start=0x7f8c0000 size=0x4000 stride=0x1000 asid=0x12 has_mm=1
 target_mask_weight=2 target_cpus=00000000,00000006 scope=range path=sbi-rfence

What This Patch Does Not Solve

Builder Takeaways

  1. TLB shootdown latency is a real portability issue on RISC-V servers. Different boards choose different paths depending on firmware and SBI version; this tracepoint would let you prove which path is taken.
  2. sbi-rfence vs cross-cpu-call is the key distinction. If firmware is slow, the SBI path will dominate; if Linux's own path is slow, you have a different tuning target.
  3. Do not script around this until v2 is accepted. The event format and even the existence of the tracepoint may change.

Related on this site: K3 PCIe Root Complex upstream v7, Linux RISC-V patch v8 for Zic64b and RVA23U64 hwprobe, KVM on the K3 reaches stock distros.

2026 年 10 月 2 日,Linux 内核邮件列表出现了 [PATCH v2] riscv: mm: Trace TLB flush path selection 的 v2 版本。补丁作者是 Roman 'Hedin' Storozhenko,它新增了一个 RISC-V 专用的 tracepoint 事件 riscv_tlb:riscv_tlb_flush_path,让运维和内核开发者能看到每一次 TLB 失效请求选择了哪条路径:本地(local)失效、通过 SBI RFENCE 调用 M-mode 固件,还是走 Linux 自己的 跨 CPU 调用(cross-cpu-call) / IPI 机制。

诚实声明。 这是评审中的 v2 补丁,尚未合入主线。10 月 2 日 Jisheng Zhang 回复,要求作者展示一个只有这个 tracepoint 才能解决的具体 bug 或性能问题,并建议用 kprobe 作为替代。事件格式与字段列表应视为暂定。

为什么 RISC-V 上的 TLB shootdown 难调试

当进程页表改变时,所有可能缓存了旧映射的 CPU 都需要让对应的 TLB 表项失效,Linux 把这叫 TLB shootdown。在 RISC-V 上,内核可以通过三种机制完成 shootdown:

路径行为内核选择条件
local仅当前 CPU 刷新自己的 TLB。请求只针对本地 CPU,或 num_online_cpus() < 2。
sbi-rfenceLinux 把远程 fence 委托给 SBI 实现(OpenSBI、厂商固件等)。riscv_use_sbi_for_rfence() 返回 true 且目标 mask 是远程的。
cross-cpu-callLinux 使用自己的跨 CPU 调用 / IPI 路径,通常是 on_each_cpu()。需要远程失效但未选择 SBI RFENCE 路径。

内核工程师面临的问题在于,此前这个选择是不可见的。一次慢的 munmap() 或 KVM VM exit 之所以慢,可能是因为 SBI 固件耗时太长、Linux 瞄准了过多 CPU,或者请求反复落在最慢路径上——但 ftrace 与 perf 只有通用的 tlb:tlb_flush 事件,它记录 reason 和 page count,不记录 RISC-V 特有的路径。

新的 tracepoint 事件

补丁在新头文件 include/trace/events/riscv_tlb.h 中定义 tracepoint 类,并在 arch/riscv/mm/tlbflush.c 的 flush_tlb_all() 与 __flush_tlb_range() 中埋点。事件记录的字段如下:

字段含义
start失效起始虚拟地址。
size失效区域大小。
stride传给 RISC-V 实现的失效粒度(不传给 SBI RFENCE)。
asid请求关联的硬件可见 ASID。
has_mm是否有特定的 mm_struct attached。
target_mask_weight目标 CPU mask 中的 CPU 数量。
target_cpus完整的目标 CPU mask,不只是数量。
scopesingle、range、address-space 或 all。
pathlocal、sbi-rfence 或 cross-cpu-call。

补丁特意记录完整 CPU mask,因为「仅靠数量无法重建 CPU 身份」——如果你想把一次慢的 shootdown 与调度活动、IPI 风暴或某个 hart 上的 SBI 调用延迟关联起来,就必须拿到 mask。

scope 与 path 语义

scope 由 size/stride 以及是否 attached mm 推导:

path 在内核即将分派失效前选定。以 flush_tlb_all() 为例,逻辑大致如下:

if (num_online_cpus() < 2)
    trace_riscv_tlb_flush_path(..., path=local);
    local_flush_tlb_all();
else if (riscv_use_sbi_for_rfence())
    trace_riscv_tlb_flush_path(..., path=sbi-rfence);
    sbi_remote_sfence_vma_asid(...);
else
    trace_riscv_tlb_flush_path(..., path=cross-cpu-call);
    on_each_cpu(__ipi_flush_tlb_all, ...);

配合 ftrace 与 perf 使用

如果补丁合入,RISC-V 板上的开发者可以用标准 trace 基础设施启用该事件:

# 启用事件
echo 1 > /sys/kernel/debug/tracing/events/riscv_tlb/riscv_tlb_flush_path/enable

# 运行触发慢 TLB shootdown 的负载
...

# 读取 trace
cat /sys/kernel/debug/tracing/trace | head -n 50

使用 perf:

perf record -e riscv_tlb:riscv_tlb_flush_path -- sleep 10
perf script

输出格式设计为可解析:

start=0x7f8c0000 size=0x4000 stride=0x1000 asid=0x12 has_mm=1
 target_mask_weight=2 target_cpus=00000000,00000006 scope=range path=sbi-rfence

这个补丁不解决什么

工程要点

  1. TLB shootdown 延迟是 RISC-V 服务器上真实的可移植性问题。 不同板卡根据固件和 SBI 版本选择不同路径;这个 tracepoint 落地后可以用来证明实际走了哪条路。
  2. sbi-rfence 与 cross-cpu-call 是关键区分。 如果固件慢,SBI 路径会占主导;如果 Linux 自己的路径慢,则调优目标不同。
  3. 在 v2 被接受前不要围绕它写脚本。 事件格式乃至 tracepoint 是否存在都可能变化。

本站相关报道:K3 PCIe Root Complex 上游 v7、Linux RISC-V v8 补丁:Zic64b 与 RVA23U64 hwprobe、K3 上的 KVM 进入发行版内核。

Краткое содержание (RU)

2 октября 2026 года Roman 'Hedin' Storozhenko опубликовал v2 патча riscv: mm: Trace TLB flush path selection. Патч добавляет tracepoint riscv_tlb:riscv_tlb_flush_path, который показывает, какой путь выбран для TLB-invalidations: local, SBI RFENCE или cross-CPU call. Событие записывает start, size, stride, ASID, target CPU mask, scope и path. Это v2, всё ещё на рассмотрении; Jisheng Zhang попросил привести реальный пример использования. Пользоваться ftrace/perf, как и любым другим tracepoint.

Resumen (ES)

El 2 de octubre de 2026, Roman 'Hedin' Storozhenko publicó la v2 del parche riscv: mm: Trace TLB flush path selection. Añade el tracepoint riscv_tlb:riscv_tlb_flush_path para observar qué camino elige el kernel para invalidaciones TLB: local, SBI RFENCE o cross-CPU call. El evento registra start, size, stride, ASID, máscara de CPUs destino, scope y path. Es v2 bajo revisión; Jisheng Zhang pidió un caso de uso real. Se puede usar con ftrace/perf.

Résumé (FR)

Le 2 octobre 2026, Roman 'Hedin' Storozhenko a publié la v2 du patch riscv: mm: Trace TLB flush path selection. Il ajoute le tracepoint riscv_tlb:riscv_tlb_flush_path pour observer quel chemin le noyau choisit pour les invalidations TLB : local, SBI RFENCE ou appel inter-CPU. L’événement enregistre start, size, stride, ASID, masque CPU cible, scope et path. Il est en v2 et en revue ; Jisheng Zhang a demandé un cas d’usage concret. Utilisable avec ftrace/perf.

Kurzfassung (DE)

Am 2. Oktober 2026 veröffentlichte Roman 'Hedin' Storozhenko v2 des Patches riscv: mm: Trace TLB flush path selection. Er fügt den Tracepoint riscv_tlb:riscv_tlb_flush_path hinzu, um zu beobachten, welchen Weg der Kernel für TLB-Invalidierungen wählt: local, SBI RFENCE oder Cross-CPU-Call. Das Ereignis zeichnet start, size, stride, ASID, Ziel-CPU-Mask, scope und path auf. Es ist v2 und unter Review; Jisheng Zhang bat um ein konkretes Anwendungsbeispiel. Verwendbar mit ftrace/perf.

خلاصه (FA)

در ۲ اکتبر ۲۰۲۶، Roman 'Hedin' Storozhenko نسخه‌ی v2 پچ riscv: mm: Trace TLB flush path selection را منتشر کرد. این پچ tracepoint riscv_tlb:riscv_tlb_flush_path را اضافه می‌کند تا مشخص شود کرده TLB کدام مسیر را انتخاب کرده: local، SBI RFENCE یا cross-CPU call. رخداد start، size، stride، ASID، target CPU mask، scope و path را شامل می‌کند. هنوز v2 است و در حال بررسی، Jisheng Zhang از متوع یک مورد کاربردی کنکرت خواست. قابل استفاده با ftrace/perf.

Sources / 参考来源

← Back to Tech Blog