RelFlexformer: Efficient Attention 3D-Transformers
for Integrable Relative Positional Encodings

Byeongchan Kim1 Arijit Sehanobish2 Avinava Dubey3
Min-hwan Oh1 Krzysztof Choromanski4,5
1Seoul National University 2Kensho Technologies 3Google Research 4Google DeepMind 5Columbia University
NeurIPS 2026
TL;DR RelFlexformer makes efficient 3D attention geometry-aware by injecting flexible relative positional encodings through NU-FFT-based FastMult, achieving spatially selective attention in \(O(L \log L)\) time.

Overview: Flexible 3D RPE with Efficient Attention

RelFlexformer keeps flexible 3D relative positional encodings compatible with Performer-style efficient attention by avoiding the explicit \(L \times L\) mask and applying it implicitly through NU-FFT-based FastMult.

RelFlexformer cartoon overview: keeping 3D RPE compatible with efficient attention

Dense Transformer vs. Performer vs. RelFlexformer (Ours)

RelFlexformer preserves the efficiency of linear attention while producing more localized, geometry-aware attention patterns in 3D scenes.

Click each panel to highlight it and read what it shows.

Input: query point

The input scene with the selected query point marked by a green star. The following panels visualize how each attention mechanism distributes attention relative to this same query location.

Why Do We Need FastMult?

Relative positional encodings can be viewed as attention masks that inject useful spatial inductive bias. However, explicitly applying a flexible 3D RPE mask requires a dense \(L \times L\) matrix, which breaks the efficiency advantage of Performer-style attention.

Masking as a powerful inductive bias in Transformers
Masking and relative positional encodings provide useful inductive biases for Transformers. The key challenge is how to keep such masks scalable when the token locations are irregular and non-uniform.
Motivation

RPE acts like a geometric mask

In 3D point clouds and scenes, token positions carry geometric meaning. A flexible RPE mask \(\mathbf{M}_{ij}=f(\mathbf{r}_i-\mathbf{r}_j)\) biases attention toward spatially relevant token interactions.

Bottleneck

Dense masks are quadratic

Naively constructing the full RPE mask requires an \(L \times L\) matrix. This introduces \(O(L^2)\) memory and computation, eliminating the scalability benefit of efficient attention.

Solution

FastMult keeps the mask implicit

RelFlexformer never materializes the dense mask. Instead, it computes the masked product \(\mathbf{M}\mathbf{u}\) directly using NU-FFT-based FastMult, enabling RPE-modulated attention in \(O(L \log L)\) time.

Key idea: keep the flexible 3D RPE mask as an inductive bias, but apply it through fast matrix-vector multiplication instead of explicitly constructing the full \(L \times L\) matrix.

NU-FFT for RPE-modulated Attention

Instead of explicitly constructing the full RPE mask, FastMult approximates the masked matrix-vector product through forward and inverse NU-FFT operations.

Algorithm 1: \(\mathrm{FastMult}_{\mathbf{M}}\): NU-FFT Mask Multiplication

Input: Input vector \(\mathbf{u} \in \mathbb{R}^{L}\), token coordinates \(\{\mathbf{r}_{i}\}_{i=1}^{L} \in \mathbb{R}^{d}\), spatial modulation function \(f\), quadrature samples \(\{\xi_{s}\}_{s=1}^{S}\) and coefficients \(\{a_{s}\}_{s=1}^{S}\).

Output: Vector \(\mathbf{w} \in \mathbb{R}^{L}\) approximating \(\mathbf{M}\mathbf{u}\).

  1. Compute the Fourier Transform of the point cloud signal at sampled frequencies \(s \in \{1, \dots, S\}\) using the forward Non-Uniform FFT in \(O(L \log L)\) time:

    \[ \mathcal{F}_{P}(\xi_{s}) = \sum_{l=1}^{L} u_{l} \exp(-2\pi i \xi_{s}^{\top} \mathbf{r}_{l}). \]
  2. Calculate the modulated coefficients for each frequency sample using the spatial mask's Fourier Transform:

    \[ b_{s} = a_{s} \mathcal{F}_{f}(\xi_{s}) \mathcal{F}_{P}(\xi_{s}). \]
  3. Output the final evaluated function at the point coordinates \(i \in \{1, \dots, L\}\) using the inverse Non-Uniform FFT in \(O(L \log L)\) time:

    \[ \mathbf{w}_{i} = \sum_{s=1}^{S} b_{s} \exp(2\pi i \xi_{s}^{\top} \mathbf{r}_{i}). \]

FastMult Analysis

Execution time as a function of sequence length and approximation accuracy of our FastMult algorithm.

Click each panel to highlight it and read what it shows.

Left: Transformer variants

Performance comparison of Transformer, Performer, and RelFlexformer models as sequence length increases. The standard Transformer runs out of memory for sequence lengths L ≥ 16k, while RelFlexformer scales to longer sequences.

Mask Behavior Analysis

RelFlexformer uses flexible spatial modulation functions to control how token interactions decay with distance. We visualize the behavior of the proposed kernels and compare them with standard Heat/RBF and Laplace kernels.

RelFlexformer mask behavior analysis
Analysis

RelFlexformer mask behavior analysis. As the Euclidean distance between points increases, the mask values decay smoothly, showing that the proposed modulation functions preserve spatial locality.

The tight scatter around the reference Heat/RBF and Laplace curves suggests stable kernel behavior across varying spatial scales.

Distance-Attention Correlation

We further analyze how attention strength changes with spatial distance from the query point. This provides a complementary view of the heatmap visualization above.

Distance-attention correlation visualization
Analysis

Distance-attention correlation. The left panel shows the spatial distance from the selected query point, while the middle and right panels visualize the corresponding attention values for Performer and RelFlexformer.

Compared with Performer, RelFlexformer shows a stronger positive relationship between spatial distance structure and attention assignment, suggesting that its RPE-modulated attention better reflects geometric organization around the query point.

Performance

RelFlexformer matches or improves strong Transformer and Performer baselines across a broad set of 3D benchmarks while operating in \(O(n \log n)\) time.


Object Classification

Attention ModelNet40 ScanObjectNN
OA mAcc OA
Transformer 93.2 80.5 84.0
Performer 92.34 80.47 83.16
+ PointRoPE 92.48 80.36 83.88
RelFlexformer (ours) 92.94 81.56 84.45
+ PointRoPE 92.55 81.49 84.26

Semantic Segmentation

Attention ScanNet Val ScanNet200 Val ScanNet++ Val nuScenes Val
mIoU mAcc allAcc mIoU mAcc allAcc mIoU mAcc allAcc mIoU mAcc allAcc
Transformer 77.6 85.0 92.0 35.3 46.0 83.4 48.2 61.6 87.0 80.4 87.2 94.7
Performer 74.8 83.8 91.0 28.2 39.6 80.2 48.1 62.5 87.2 72.0 80.2 93.5
+ PointRoPE 74.6 83.4 91.0 34.1 44.9 82.9 48.1 60.9 86.9 80.4 87.4 94.8
RelFlexformer (ours) 76.8 85.0 91.9 34.0 45.4 82.9 48.7 62.2 86.8 80.3 87.5 94.6
+ PointRoPE 76.6 84.5 91.6 34.9 45.2 82.9 48.8 61.7 86.7 81.2 87.5 94.8

Indoor Semantic Segmentation

Attention Metric Area1 Area2 Area3 Area4 Area5 Area6 6-Fold
Transformer allAcc 93.22 86.26 94.56 90.72 91.67 94.98 91.90
mAcc 89.92 74.44 94.45 81.11 78.92 93.55 85.31
mIoU 83.01 63.42 86.66 71.34 73.43 87.31 77.70
Performer allAcc 92.35 88.53 94.47 85.96 90.86 91.20 90.56
mAcc 88.56 76.97 93.69 78.31 75.63 85.77 83.16
mIoU 80.74 62.60 86.02 62.60 69.75 77.35 73.18
+ PointRoPE allAcc 93.00 86.88 94.25 88.29 91.17 94.73 91.39
mAcc 90.06 76.62 93.63 76.14 77.55 93.26 84.54
mIoU 82.31 61.47 86.21 67.38 71.75 86.59 75.95
RelFlexformer (ours) allAcc 92.96 88.07 94.51 88.36 91.14 94.65 91.62
mAcc 90.59 77.33 93.00 80.92 76.79 93.33 85.33
mIoU 81.93 63.21 86.20 68.28 71.01 86.88 76.25
+ PointRoPE allAcc 93.32 85.84 94.57 88.89 91.50 94.74 91.48
mAcc 90.43 73.82 93.64 81.59 77.44 93.14 85.01
mIoU 82.79 62.87 86.75 68.12 72.14 86.79 76.58

RGB-D Semantic Segmentation

Attention Complexity NYU v2 SUN RGB-D
Softmax \(O(n^2)\) 55.6 51.2
Performer \(O(n)\) 54.44 48.49
Performer + STRING \(O(n \log n)\) 54.96 50.87
RelFlexformer (ours) \(O(n \log n)\) 55.32 51.04

Citation

@article{kim2026relflexformer,
  title={RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings},
  author={Kim, Byeongchan and Sehanobish, Arijit and Dubey, Avinava and Oh, Min-hwan and Choromanski, Krzysztof},
  journal={arXiv preprint arXiv:2605.10706},
  year={2026}
}