{\Phi}eat: Physically Grounded Material Feature Representation
About
While foundation models have emerged as general-purpose visual backbones, their representations are primarily optimized for semantics and lack explicit modeling of physical factors, such as reflectance, hindering their efficacy in tasks requiring explicit material reasoning. We introduce $\Phi$eat$, a novel material-grounded visual backbone that encourages a representation sensitive to material identity, including reflectance and mesostructure. Instead of relying on generic data augmentations, we pretrain our model by contrasting observations of the same material under controlled variations in lighting and geometry. This encourages invariance to extrinsic factors while preserving sensitivity to intrinsic material properties. We show that the resulting representation provides strong priors for material-centric tasks, including feature-based material selection and classification. Our results demonstrate that physically inspired weak supervision is an effective strategy for learning representations tailored to material perception.
Related benchmarks
| Task | Dataset | Result | Rank | |
|---|---|---|---|---|
| Material Selection | DuMaS | L1 Error25.4 | 10 | |
| k-NN classification | Synthetic (test) | Accuracy71.8 | 8 | |
| Material Consistency | Material renderings Illumination variations | Hamming Distance0.221 | 4 | |
| Material Consistency | Material renderings Geometry variations | Hamming Distance30.5 | 4 |