de en es fr nl pl pt sv zh

research paper

Speculative Decoding in LLM Inference

Jeffrey Wang

Running frontier LLMs is slow, but that's the tradeoff we make to get more intelligent output. But if you think about the text (tokens) that an LLM produces (or that a human produces, for that matter), you might have an intuition that a lot of it does not actually require high intelligence....

Triton Language

Jeffrey Wang

The world of GPU programming for AI has come a long way since I worked on writing a CUDA-based matrix library back in 2009. Both NVIDIA hardware and the CUDA ecosystem have evolved dramatically and are now the basis for the majority of AI compute in the world today (hence the $4T+ market cap)....

A Vision for the Team: The Conclusion of my Series on Design Team Strategy

Maximilian Speicher

Tales of Design & User Experience (ToDUX) #13 ⁂ Dear readers & friends, It’s been a while, and I’m sorry for that. I’ve been suffering from a mild case of writer’s block for the better part of the past year. But I’ve managed to turn things around and could finally finish my...