

9 hours ago
The following is compiled by Dolphin Research for $MINIMAX-W(00100.HK) FY26 interim earnings call Trans
I. Key takeaways from the results
1) Forward-looking commentary (no formal revenue guidance this call; only the following outlook).
a) GPM will continue to improve in H2 and further beyond 2026.
b) Domestic chips will account for a larger share across models in Q4, which management believes will lower per-token costs.
c) As the H and M pipelines mature, release cycles for new models will shorten materially.
2) Revenue and mix: H1 revenue of $112 mn (+283% YoY), already 1.5x FY25 full-year revenue.
a) Open platform and other AI enterprise services revenue was approx. $74 mn (+703% YoY), with mix rising to 63% vs. 30% a year ago, driven by paid user and enterprise client growth, higher API usage, and rapid adoption of token bundles.
b) AI-native product revenue was approx. $43 mn (+101% YoY). Overseas contributed over 60% of total revenue.
3) GP and opex.
a) GP reached $21 mn (+465% YoY); GPM expanded 580 bps to 17.9% from 12.1%.
b) R&D expense was approx. $300 mn (+138% YoY), mainly into foundation models; management stressed R&D growth lagged the 283% revenue growth.
c) S&D expense was approx. $30 mn (-18% YoY) on organic user growth and lower promo spend. G&A was $30 mn (+104% YoY), with revenue ratio down to 26% from 48%.
4) Loss and cash.
a) IFRS net loss narrowed 11% YoY to approx. $400 mn.
b) Ex-SBC, FV loss on financial liabilities and listing fees, Adj. net loss was approx. $290 mn vs. approx. $139 mn a year ago, widening 112% YoY—opposite to the IFRS trend.
c) A placement in Jul bolstered cash and financial flexibility; cash and cash equivalents exceeded $3 bn at period-end.
II. Details from the earnings call
2.1 Management remarks
1) Biz. progress and operating metrics.
a) Continued upgrades to the M2 series this year, and launched M3 and H3. With higher model intelligence and inference throughput, the overall biz. has scaled rapidly.
b) Jul token usage reached 20x Jan levels, with revenue up 81.8% vs. Jan. ARR rose to $800 mn in Aug.
2) Strategy: minimize unit intelligence cost.
a) Intelligence gains from LLMs are far from capped. Since H2 last year, coding and agent capabilities became practical, with models moving toward higher autonomy and creativity for long-horizon tasks, while energy and compute remain finite.
b) The vision is 'Intelligence with Everyone'—achieved by minimizing cost while maximizing intelligence. Management stressed this is not merely value pricing; it is a core capability required to scale toward higher-order intelligence.
3) Pre-training and architectural efficiency.
a) Exploring two extremes on large-parameter architectures: on the sparse end, keeping activation within 2% while maintaining good convergence; and combining dense with sparse attention.
b) Introduced MiniMax sparse attention to keep training stable while reducing compute needs for long-context processing.
c) For the flagship model in development, a new MSAR architecture lowers KV cache VRAM demand, raises pre-fill cache hit rates, and supports larger batch sizes to improve compute utilization.
d) At max parameter scale, compute efficiency is 3x the first generation, with bigger advantages on longer-context data.
4) Infra and full-stack capability.
a) One of the earliest independent foundation model firms in China to build scalable, stable, long-term infra, and among the first two to lock in long-term capacity. Self-built training infra achieved 97% ETTR, the industry’s highest.
b) Through co-design of memory, SSD, and network, KV cache hit rates improved significantly, sharply reducing pre-fill costs.
c) The multimodal inference cluster includes ample CPU resources to support post-training CPU sandboxes and heterogeneous scheduling. Traffic peaks and troughs in online inference free capacity for small validation experiments and RL rollouts.
d) With full-stack capability, management believes post-training scale per unit capital can still grow at least 3x via inference efficiency, while supporting higher GPM.
5) Multimodal and self-reflection.
a) Unlike many LLM companies, the models natively handle multimodal inputs from the start, as visual understanding and generation are core to productivity. Multimodal generation, alongside coding, is another major AGI market.
b) H3 is the first integrated effort combining language and visual generation, using the language model not just as an encoder but for fine-grained context understanding to fully leverage reference inputs and enable fine-grained generation.
c) H3’s performance breakthrough, open-source strategy, and value reshaped a market previously dominated by big tech’s closed models. Downloads topped 24 mn within three weeks of open-sourcing, with hosted usage surging and community response far above expectations.
d) Management acknowledged shortcomings in judgment and execution during M3 development. Any single-model lead is not durable; the goal is to continuously define and refine the technical roadmap.
2.2 Q&A
Q: What underpins the view that intelligence can keep improving? Beyond coding, which use cases will drive scale?
A: Technically, long-horizon task capability is deliverable once verifiable; coding still isn’t end-to-end. In recent months, the yardstick has been handling long-horizon tasks—often requiring hundreds of tool calls and architecture checks. Internal and external tests show that once a capability is measurable or verifiable, the model can deliver it reliably, which is highly certain.The key variable is abstracting these tasks into specific environments for the model to learn. On top of sufficient pre-training, SFT and RL drive better performance, with higher-quality data in RL yielding stronger capabilities. Looking ahead, RL scale will determine iteration speed, and we are investing heavily here.On the market side, coding is in demand but still early—largely code completion and generation, not end-to-end delivery. Within broad coding, cybersecurity is a clear example: once model coding ability is strong enough, it can deliver exponential improvement in cyber defense.
As long as we can build closed loops and establish reputation, models will penetrate verticals such as chip design, with more such scenarios emerging in the coming months.
Q: Can you break down token consumption and ARR growth—new users, existing users ramping, or product diversification? Any regional color?
A: Both new users and per-user usage rose; B2B is 80% of ARR, overseas ~60%. Token consumption surged on two fronts: agentic-related usage and API growth driven by M3’s native capabilities, with M3 being more compute-efficient; together they lifted consumption.ARR growth is mainly from model capability gains: enterprise users topped 2 mn, 10x YoY, with existing users unlocking new office scenarios. Consumption is shifting from human-to-model toward model-to-agent, where multi-turn requests and tool calls grow much faster than human interactions, driving rapid per-user usage growth.
In Aug, B2B accounted for 80% of ARR vs. ~30% a year ago. In H1, overseas revenue was ~60% of total. Management views ‘the model is the product’—whether 2B or 2C, customers pay for model capability.By modality, multimodal usage rose sharply after H3, while text models grew on value, with more users adopting both. From Jun to Aug, both contributed to ARR growth. In H2, the main drivers remain capability gains and higher inference efficiency.
Q: Peers are raising or cutting prices. How is ‘lowering inference cost’ not just a commercial strategy but also a way to raise model intelligence?
A: Inference itself produces intelligence; inference efficiency sets the scaling law ceiling. During pre-training, parameters scale and the scaling law holds; post-training, inference takes a rising share.Post-training relies heavily on synthetic data, agentic environments, and sampling. RL involves massive rollouts, numerous experiments, uncertainties, and precise evaluations—much of the compute is inference load.Thus, inference efficiency is key to the scaling law’s ceiling: holding other factors equal, the higher the inference efficiency, the more gains we can extract from post-training. Efficiency is not just cost—it's also how many inference cycles, experiments, and evaluations can run per unit time.
Externally it shows up as pricing power; internally it dictates iteration speed. This is where we see competitive advantage for the next generation of models.
Q: What is the timeline and cadence for M3.1 and M3 Pro? In coding agents and long-horizon tasks, which scenarios can validate your products?
A: No timeline given; M3 Pro is a 3-trillion-parameter model. Releases are only one piece; the crux is reusable infra—compute, data and evaluation frameworks, inference infra, end-to-end supply chain optimization, and cross-team coordination between infra and algorithms.We faced challenges during M3, especially on evaluation infra, which is why we are rebuilding much of it. As this infra matures, we are confident about the next-gen products.
M3.1 aims to flatten and close the loop on the underlying infra, with comprehensive evaluation across stability and generalization; meeting expected results is the goal. M3 Pro has 3 tn parameters, pursuing the scaling law’s limits across scale, training efficiency, and every step of post-training, with architectural innovation to lift inference efficiency and speed in tandem.Exact magnitudes are being assessed, but management is confident M3 Pro will lead peers. H3 has seen broad adoption with open-source support, and the next generation will raise the bar again.
Q: Do you have sufficient compute reserved for M3 and future models? How do you split training vs. inference? How compatible are domestic chips?
A: Compute is adequate for the pipeline; domestic chip share will rise in Q4. Leveraging self-build experience, we are reserving more compute for future iterations; capacity is sufficient for the 3-tn-parameter model and video generation. Compute sources are threefold: self-owned, hyperscalers, and token factories, forming a complete network.Self-owned compute is for flagship training: large clusters must run continuously and require high HW/SW co-optimization and high-bandwidth networks and private lines—investments that will show up in financials. We made clear progress in H1, laying groundwork for rapid iterations.
Inference loads go mainly to hyperscalers. We also partner with token factories as an alternative source to add capacity per actual token needs and absorb spikes, ensuring service quality while helping token factories scale.Model needs vary by stage, and internal resource orchestration is a core strength. Text is at a critical stage, with training load well above video generation, and LLMs underpin other multimodal products. We are already using some domestic chips, with their share rising in Q4.
Q: With token prices trending down, how will you sustain the sharp GPM improvement?
A: Unit compute throughput is up 3x; price cuts and GPM gains are not contradictory. Margin gains come from supply chain, management, and operating leverage, with ample room to improve.First and foremost is compute resource optimization: unit compute throughput rose 3x; price cuts offset part of the gains, but improvement remains substantial. Second is training–inference coordination: when nighttime inference requests drop, we shift compute to evaluation and algorithm checks to lift overall utilization; self-built infra gives greater scheduling flexibility. Third is supply chain, which still has room for optimization as self-build scales.
Management sees no conflict between price cuts and GPM expansion: as costs fall, lowering prices broadens access, and scale then drives further GPM gains. Once the flywheel turns, larger scale means more margin upside.By mix, multimodal and audio generation carry higher GPM. Text models are driving rapid revenue growth and, with further cluster optimization, will be key to margin expansion. GPM will continue to improve in H2 and further beyond 2026.
Q: Regarding H3’s launch and open-sourcing—how do you validate the multimodal path? Will open-source weaken pricing and monetization?
A: H3’s reception validates the multimodal decision; open-source enhances, not weakens, monetization. For a long time, investors repeatedly asked whether to go multimodal; H3’s reception affirmed that choice.We chose open-source because multimodality is tightly linked to productivity, and our goal with H3 was to change a landscape dominated by closed, high-priced models. As with language models, we believe healthy cycles are open.
We are still early: while the community is enthusiastic and H3 performs well, reliability, output quality, and source trustworthiness need work, with each iteration moving the industry forward. Open-source expands our room to monetize these capabilities, and H3 is a strong start.
Q: How do you balance resources while advancing both text and multimodal models?
A: We do not pit text against multimodal; they advance in tandem. They reinforce each other at the base layer, and we pursue synergy to create stronger combined effects. The goal is to build productivity tools and broaden access.On text vs. video generation, we expect more frontier fields such as chip design and drug discovery, where multiple modalities can be fused—an already proven path. By Q4 or Q1 next year, more opportunities should emerge, deepening the fusion between text and video generation and between text and multimodality.
Q: Competition is intensifying—what is management’s view?
A: We do not see a zero-sum; the industry is still at the start, and iteration speed is key. No quantified share or peer comparisons were provided. Qualitatively, over the next 3–6 months, deployable intelligence remains early, and any firm that advances the frontier will enlarge the market and benefit the ecosystem.
Q: What will be MiniMax’s long-term moat?
A: The key is not sheer compute, but the ability to frame problems and choose effective paths. There are many choices across architecture, compute, RL, and evaluation, and different companies will pursue very different approaches, leading to divergent solutions.Delivering the best value while raising intelligence is not just pricing; it is a technical strategy. As AI boosts productivity, models will still be judged by real-world task performance, and users have shown they will pay for better-performing models.
We do not frame it as startups vs. big tech. The crux is who can reach higher intelligence, lower unit intelligence cost, and raise conversion efficiency.
<End of text>
Risk disclosure and statement:Dolphin Research disclaimer and general disclosure
Login to unlock14,748characters for free
This content is only available to signed-in users. Sign in to your Longbridge account to read the full post.
