Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
YetAnotherNick
37 days ago
|
parent
|
context
|
favorite
| on:
Accelerating GPT-5.6 Sol Ultrafast
LLama 3 405B had the most unoptimized kv cache usage by far. Deepseek v4 pro uses 2.4GB for the same context length[1].
[1]:
https://vllm.ai/blog/2026-04-24-deepseek-v4
philipportner
37 days ago
[–]
Good point, thanks! I haven't been keeping up with most of the new model internals.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search:
[1]: https://vllm.ai/blog/2026-04-24-deepseek-v4