Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

Implicit Neural Representations of Individual Behavior

About

We study policy representation learning from unlabeled multi-policy behavioral data. Each episode is generated by a fixed policy, but policy labels are unavailable. This setting appears in robotics play, demonstrations, games, racing, and other datasets where heterogeneous behaviors are mixed without annotations. We introduce \emph{Behavioral INR}, a self-supervised generative model that adapts implicit neural representations (INRs) from vision to behavior. Instead of mapping coordinates to RGB values, Behavioral INR represents a policy as a state-action function mapping states to subsequent actions. An episode-level latent modulates this function through FiLM layers, yielding a generative prior over policies and allowing policy identity to be inferred without supervision. Because INRs treat each datapoint as samples from an underlying function, the same model naturally accommodates variable episode lengths and different sampling granularities, as in vision INRs with different image resolutions. We also define policy-level out-of-distribution (OOD) shifts along state-distribution and action-distribution axes, which arise when policies overlap in states or actions but are not captured by standard behavioral OOD settings based only on new agents or environments. We evaluate on synthetic Gaussian random field data, MuJoCo demonstrations with controlled OOD splits, and real-world chess, Formula 1 racing, robotics, and Seek-Avoid datasets. Behavioral INR most consistently improves policy identifiability in the hardest continuous state-action settings, especially when longer episodes, more policies, and OOD splits reduce the usefulness of marginal shortcuts; amortized history encoders remain competitive when policy identity can be recovered from symbolic repetition or low-dimensional action statistics. We release code and checkpoints.

Andrew Kang, Priya Narasimhan• 2026

Related benchmarks

TaskDatasetResultRank
Policy IdentificationFastF1 SPECIALIZATION
Probe Accuracy19
7
Policy Identity RecoverySynthetic GRF 10x all-policy (GENERALIZATION)
Probe Accuracy61.1
7
Policy IdentificationLichess SPECIALIZATION
Probe Accuracy50.24
7
Policy IdentificationDROID Specialization
Probe Accuracy50
7
Behavior Representation GeneralizationSynthetic GRF 10x all policies (aggregate)
P Score61.1
4
Continuous Action Prediction GeneralizationHopper 20x (aggregate)
P0.744
4
Behavior Representation LearningDM Lab Seek-Avoid Specialization
Probe Accuracy100
4
Behavior Representation GeneralizationDROID all policies (aggregate)
P Score0.444
4
Behavior Representation GeneralizationFastF1 all drivers (aggregate)
P Metric Value5.3
4
Behavior Representation GeneralizationLichess all policies (aggregate)
P Score0.33
4
Showing 10 of 11 rows

Other info

Follow for update