Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

TriLift: Interpolation-Free Tri-Plane Lifting for Efficient 3D Perception on Embedded Systems

About

Dense 3D convolutions provide high accuracy for perception but are too computationally expensive for real-time robotic systems. Existing tri-plane methods rely on 2D image features with interpolation, point-wise queries, and implicit MLPs, which makes them computationally heavy and unsuitable for embedded 3D inference. As an alternative, we propose TriLift, a novel interpolation-free tri-plane lifting and volumetric fusion framework that directly projects 3D voxels into plane features and reconstructs a feature volume through broadcast and summation. This shifts nonlinearity to 2D convolutions, reducing complexity while remaining fully parallelizable. To mitigate spatial information loss inherent in projections, we incorporate a lightweight adaptive positional encoding module that helps bridge the spatial information gap, dynamically recovering fine geometric details with negligible overhead. To capture global context, we add a low-resolution volumetric branch fused with the lifted features through a lightweight integration layer, yielding a design that is both efficient and end-to-end GPU-accelerated. To validate the effectiveness of the proposed method, we conduct experiments on classification, completion, segmentation, and detection, and we map the trade-off between efficiency and accuracy across tasks. Results show that classification and completion retain or improve accuracy, while segmentation and detection show a trade-off, significantly reducing computational demand with only a slight decrease in accuracy. On-device benchmarks on an NVIDIA Jetson Orin Nano confirm robust real-time throughput, demonstrating the suitability of the approach for embedded robotic perception.

Sibaek Lee, Jiung Yeon, Hyeonwoo Yu• 2025

Related benchmarks

TaskDatasetResultRank
3D Shape CompletionModelNet40
Chamfer Distance (L2)0.0014
5
3D Object DetectionHypersim
Recall @ 25%60.32
5
3D Object DetectionScanNet
Recall @ IoU=0.2588.67
5
3D ClassificationModelNet40
Accuracy83.54
5
3D Object Detection3D-FRONT
R@2597.06
5
ClassificationModelNet
FPS49.58
4
CompletionModelNet
FPS35.79
4
Object DetectionScanNet
FPS10.64
4
Semantic segmentationScanNet
FPS10.87
4
Showing 9 of 9 rows

Other info

Follow for update