Skip to main content
SandRise logo SandRise
Exploring Next / Topics / Speculative Decoding

Topic

Speculative Decoding

3 episodes

  1. Ep 797 Jul 28, 2026

    Kimi K3 Is Here: Efficient Day 0 Support on vLLM

    Vince and Ava unpack Moonshot AI's Kimi K3, a 2.8‑trillion‑parameter multimodal MoE, and its day‑zero support in vLLM. They walk through the model’s hybrid attention, the engineering tricks that make a 1 M‑token context feasible, the practical deployment recipe, and how it stacks up against other frontier models.

    New ModelsInferenceLaunchKimi K3
  2. Ep 777 Jul 24, 2026

    Overview: Decoding Strategy

    We finally slow down on decoding strategy, the rule that turns a model's next-token odds into the actual words you see. We use one hallway-and-doors picture to make greedy decoding, sampling, top-k, top-p, beam search, and newer decoding work feel less like magic knobs.

    InferenceHugging Face TransformersOpenAI CodexN T T Data
  3. Ep 696 Jul 17, 2026

    Exploring Next Overview: Speculative Decoding

    We finally slow down and unpack speculative decoding from the ground up: the draft model, the verify step, and why it can make generation faster without changing the output. We keep it concrete, because that trick sounds like cheating until the mechanism actually clicks.

    InferenceVllmSglangTensorrt LLM
SandRise logo SandRise Product Studio
Resume LinkedIn GitHub Email

© 2026 SandRise · Built by Nick Sanders

🧠 PM Perspective

Crafting your PM challenge
Analyzing context and generating a thoughtful question...
Your Challenge
0 / 2000
✨

Feedback on Your Answer

⚠️

Say Hi

Feedback, ideas, interesting finds — anything goes.

What's this about?
0 / 2,000

Note received!

Thanks for reaching out. I'll take a look soon.

⚠️

Something went wrong. Please try again.