Currently working on novel hardware-aware pretraining architectures for low-latency inference at scale, diffusion language models, sample-efficient context extension, and spectral clipping optimizers.
Feel free to email me — I am almost always interested in meeting new people.