- DeepSeek is advancing its new version model, with a greyscale test before the official V4 launch, featuring an expanded context length from 128K to 1M.
- Nomura Securities reports that the upcoming V4 model aims to drive AI commercialization through innovative architecture, significantly lowering training and inference costs.
- This shift is expected to stimulate demand for AI applications and reshape the competitive landscape, benefiting Chinese AI hardware manufacturers while software companies might see enhanced value.
- DeepSeek's research presents a new "Engram" module that breaks the expensive paradox of the Transformer architecture by separating memory from computation in AI models.
- This separation allows static knowledge to be stored efficiently while freeing up model capacity for more complex reasoning tasks, leading to unexpected improvements in logic and coding capabilities.
- The upcoming DeepSeek V4 model, set to release before the Lunar New Year, aims to capitalize on these advancements, potentially redefining AI standards by enhancing memory capacity and inference efficiency.