← All guests — NVIDIA / DeepSeek V4.1 Flash
Chris Alexiuk
NVIDIA · DeepSeek V4.1 Flash
Chris Alexiuk joined Alex Volkov on ThursdAI to break down DeepSeek V4.1 Flash. The NVIDIA product research engineer walked through what makes the release interesting under the hood: the disaggregated serving design DeepSeek baked into the model rather than bolting on after the fact, and why the KV cache is the quiet secret behind what inference actually costs.
Full segment
Chris on ThursdAI
Bangers
They baked disaggregation in
DeepSeek didn't bolt disaggregated serving on afterwards — V4.1 Flash was designed with it baked in.
KV cache is the cost secret
The KV cache is the quiet lever behind what serving a model like V4.1 Flash actually costs.
Share it
Suggested posts
𝕏 Post
DeepSeek V4.1 Flash is a really interesting release once you look under the hood. I joined @altryne on ThursdAI to break it down — the disaggregated serving design they baked in, and why the KV cache is the real story on inference cost. Full segment:
https://thursdai.news/guests/llm_wizard/sep-10-2026
LinkedIn
I joined Alex Volkov on ThursdAI this week to break down DeepSeek V4.1 Flash.
We went under the hood of the release: the disaggregated serving design DeepSeek baked into the model rather than bolting on afterwards, and why the KV cache is the quiet lever behind what inference actually costs.
Full segment and clips: https://thursdai.news/guests/llm_wizard/sep-10-2026
ThursdAI goes live every Thursday and is one of the highest signal AI shows around. Worth a follow.
ThursdAI — The weekly AI podcast, hosted by Alex Volkov. Every Thursday, live.
Subscribe Free →