Object Classification
| Attention | ModelNet40 | ScanObjectNN | |
|---|---|---|---|
| OA | mAcc | OA | |
| Transformer | 93.2 | 80.5 | 84.0 |
| Performer | 92.34 | 80.47 | 83.16 |
| + PointRoPE | 92.48 | 80.36 | 83.88 |
| RelFlexformer (ours) | 92.94 | 81.56 | 84.45 |
| + PointRoPE | 92.55 | 81.49 | 84.26 |
RelFlexformer keeps flexible 3D relative positional encodings compatible with Performer-style efficient attention by avoiding the explicit \(L \times L\) mask and applying it implicitly through NU-FFT-based FastMult.
RelFlexformer preserves the efficiency of linear attention while producing more localized, geometry-aware attention patterns in 3D scenes.
The input scene with the selected query point marked by a green star. The following panels visualize how each attention mechanism distributes attention relative to this same query location.
Relative positional encodings can be viewed as attention masks that inject useful spatial inductive bias. However, explicitly applying a flexible 3D RPE mask requires a dense \(L \times L\) matrix, which breaks the efficiency advantage of Performer-style attention.
In 3D point clouds and scenes, token positions carry geometric meaning. A flexible RPE mask \(\mathbf{M}_{ij}=f(\mathbf{r}_i-\mathbf{r}_j)\) biases attention toward spatially relevant token interactions.
Naively constructing the full RPE mask requires an \(L \times L\) matrix. This introduces \(O(L^2)\) memory and computation, eliminating the scalability benefit of efficient attention.
RelFlexformer never materializes the dense mask. Instead, it computes the masked product \(\mathbf{M}\mathbf{u}\) directly using NU-FFT-based FastMult, enabling RPE-modulated attention in \(O(L \log L)\) time.
Instead of explicitly constructing the full RPE mask, FastMult approximates the masked matrix-vector product through forward and inverse NU-FFT operations.
Input: Input vector \(\mathbf{u} \in \mathbb{R}^{L}\), token coordinates \(\{\mathbf{r}_{i}\}_{i=1}^{L} \in \mathbb{R}^{d}\), spatial modulation function \(f\), quadrature samples \(\{\xi_{s}\}_{s=1}^{S}\) and coefficients \(\{a_{s}\}_{s=1}^{S}\).
Output: Vector \(\mathbf{w} \in \mathbb{R}^{L}\) approximating \(\mathbf{M}\mathbf{u}\).
Compute the Fourier Transform of the point cloud signal at sampled frequencies \(s \in \{1, \dots, S\}\) using the forward Non-Uniform FFT in \(O(L \log L)\) time:
Calculate the modulated coefficients for each frequency sample using the spatial mask's Fourier Transform:
Output the final evaluated function at the point coordinates \(i \in \{1, \dots, L\}\) using the inverse Non-Uniform FFT in \(O(L \log L)\) time:
Execution time as a function of sequence length and approximation accuracy of our FastMult algorithm.
Performance comparison of Transformer, Performer, and RelFlexformer models as sequence length increases. The standard Transformer runs out of memory for sequence lengths L ≥ 16k, while RelFlexformer scales to longer sequences.
RelFlexformer uses flexible spatial modulation functions to control how token interactions decay with distance. We visualize the behavior of the proposed kernels and compare them with standard Heat/RBF and Laplace kernels.
RelFlexformer mask behavior analysis. As the Euclidean distance between points increases, the mask values decay smoothly, showing that the proposed modulation functions preserve spatial locality.
The tight scatter around the reference Heat/RBF and Laplace curves suggests stable kernel behavior across varying spatial scales.
We further analyze how attention strength changes with spatial distance from the query point. This provides a complementary view of the heatmap visualization above.
Distance-attention correlation. The left panel shows the spatial distance from the selected query point, while the middle and right panels visualize the corresponding attention values for Performer and RelFlexformer.
Compared with Performer, RelFlexformer shows a stronger positive relationship between spatial distance structure and attention assignment, suggesting that its RPE-modulated attention better reflects geometric organization around the query point.
RelFlexformer matches or improves strong Transformer and Performer baselines across a broad set of 3D benchmarks while operating in \(O(n \log n)\) time.
| Attention | ModelNet40 | ScanObjectNN | |
|---|---|---|---|
| OA | mAcc | OA | |
| Transformer | 93.2 | 80.5 | 84.0 |
| Performer | 92.34 | 80.47 | 83.16 |
| + PointRoPE | 92.48 | 80.36 | 83.88 |
| RelFlexformer (ours) | 92.94 | 81.56 | 84.45 |
| + PointRoPE | 92.55 | 81.49 | 84.26 |
| Attention | ScanNet Val | ScanNet200 Val | ScanNet++ Val | nuScenes Val | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| mIoU | mAcc | allAcc | mIoU | mAcc | allAcc | mIoU | mAcc | allAcc | mIoU | mAcc | allAcc | |
| Transformer | 77.6 | 85.0 | 92.0 | 35.3 | 46.0 | 83.4 | 48.2 | 61.6 | 87.0 | 80.4 | 87.2 | 94.7 |
| Performer | 74.8 | 83.8 | 91.0 | 28.2 | 39.6 | 80.2 | 48.1 | 62.5 | 87.2 | 72.0 | 80.2 | 93.5 |
| + PointRoPE | 74.6 | 83.4 | 91.0 | 34.1 | 44.9 | 82.9 | 48.1 | 60.9 | 86.9 | 80.4 | 87.4 | 94.8 |
| RelFlexformer (ours) | 76.8 | 85.0 | 91.9 | 34.0 | 45.4 | 82.9 | 48.7 | 62.2 | 86.8 | 80.3 | 87.5 | 94.6 |
| + PointRoPE | 76.6 | 84.5 | 91.6 | 34.9 | 45.2 | 82.9 | 48.8 | 61.7 | 86.7 | 81.2 | 87.5 | 94.8 |
| Attention | Metric | Area1 | Area2 | Area3 | Area4 | Area5 | Area6 | 6-Fold |
|---|---|---|---|---|---|---|---|---|
| Transformer | allAcc | 93.22 | 86.26 | 94.56 | 90.72 | 91.67 | 94.98 | 91.90 |
| mAcc | 89.92 | 74.44 | 94.45 | 81.11 | 78.92 | 93.55 | 85.31 | |
| mIoU | 83.01 | 63.42 | 86.66 | 71.34 | 73.43 | 87.31 | 77.70 | |
| Performer | allAcc | 92.35 | 88.53 | 94.47 | 85.96 | 90.86 | 91.20 | 90.56 |
| mAcc | 88.56 | 76.97 | 93.69 | 78.31 | 75.63 | 85.77 | 83.16 | |
| mIoU | 80.74 | 62.60 | 86.02 | 62.60 | 69.75 | 77.35 | 73.18 | |
| + PointRoPE | allAcc | 93.00 | 86.88 | 94.25 | 88.29 | 91.17 | 94.73 | 91.39 |
| mAcc | 90.06 | 76.62 | 93.63 | 76.14 | 77.55 | 93.26 | 84.54 | |
| mIoU | 82.31 | 61.47 | 86.21 | 67.38 | 71.75 | 86.59 | 75.95 | |
| RelFlexformer (ours) | allAcc | 92.96 | 88.07 | 94.51 | 88.36 | 91.14 | 94.65 | 91.62 |
| mAcc | 90.59 | 77.33 | 93.00 | 80.92 | 76.79 | 93.33 | 85.33 | |
| mIoU | 81.93 | 63.21 | 86.20 | 68.28 | 71.01 | 86.88 | 76.25 | |
| + PointRoPE | allAcc | 93.32 | 85.84 | 94.57 | 88.89 | 91.50 | 94.74 | 91.48 |
| mAcc | 90.43 | 73.82 | 93.64 | 81.59 | 77.44 | 93.14 | 85.01 | |
| mIoU | 82.79 | 62.87 | 86.75 | 68.12 | 72.14 | 86.79 | 76.58 |
| Attention | Complexity | NYU v2 | SUN RGB-D |
|---|---|---|---|
| Softmax | \(O(n^2)\) | 55.6 | 51.2 |
| Performer | \(O(n)\) | 54.44 | 48.49 |
| Performer + STRING | \(O(n \log n)\) | 54.96 | 50.87 |
| RelFlexformer (ours) | \(O(n \log n)\) | 55.32 | 51.04 |
@article{kim2026relflexformer,
title={RelFlexformer: Efficient Attention 3D-Transformers for Integrable Relative Positional Encodings},
author={Kim, Byeongchan and Sehanobish, Arijit and Dubey, Avinava and Oh, Min-hwan and Choromanski, Krzysztof},
journal={arXiv preprint arXiv:2605.10706},
year={2026}
}