How we shipped a live sports AI companion to 1.1M concurrent viewers — scaling speech-to-speech voice agents from prototype to production.
Why matrix multiplication makes GPUs essential for AI.
A walkthrough of LUMOS — a single transformer that replaces task-specific user-behavior models. Trained on 1.7T tokens from 250M users; +3.15% DAU in production A/B test.
Invited talk on building and deploying the LUMOS foundation model.
Reviewing three papers from DeepMind, Google Research, and Sakana AI that signal a shift from post-training and RL to test-time scaling and evolutionary algorithms.
Presenting the ForeCal paper on DNN calibration.
Fundamentals of credit risk modeling and practical approaches.