Abstract
NdLinear applies separate transformations along tensor dimensions instead of flattening every input into a single vector. The paper studies parameter efficiency and applications across model families, including low-rank adaptation, while examining limitations when interactions are not separable across dimensions.
Publication
arXiv preprint

AI Researcher
AI researcher at Meta MSL with a background in information theory and computational neuroscience, working on world models, memory, and compression. Former Assistant Professor and Faculty Fellow at NYU. My industry experience spans Wand AI, Intel, Google AI, and Wikipedia.