AI researcher. Currently doing a Ph.D. on multimodal large language models. Small models, big ideas, and the careful work of grounding language in pixels.
How 4-bit floats work, why MXFP4 loses small values, what NVFP4 changes about scaling, and the four-part recipe NVIDIA used to pretrain a 12B model in NVFP4 at FP8 accuracy.
Read the bit →