

6 hours ago
Dolphin Research summary of $Z.AI(02513.HK) FY26 interim earnings call
I. Key takeaways from the results
1. First-time ARR disclosure with two methodologies
a. As of end-Aug 2026, ARR reached $1.6 bn using a monthly annualized method. This equals Aug revenue multiplied by 12.
b. Management also referenced a more aggressive industry method, annualizing the latest week by 52. Under this approach, and factoring in GLM-5.3 ramp, the latest ARR is above $2.0 bn.
2. Revenue mix shifted with both volume and price up
a. Total: 1H revenue was RMB 957 mn (about $142 mn), up nearly 400% YoY. Growth was broad-based.
b. By segment: Open platform and API revenue was RMB 825 mn (about $123 mn), up over 27x YoY. Its mix rose from 15.2% a year ago and 26.3% at year-end to 86.5%.
c. Volume vs. price: Avg. API pricing rose about 101% while coding plan calls grew more than 23x vs. early-year levels. Management stressed the ramp is driven by model capability rather than price elasticity.
3. GPM and compute efficiency improved in tandem
a. Open platform and API GPM improved from -0.4% to 24.6% YoY, an expansion of 25ppt. It also rose vs. the FY level of 18.9% by around 6ppt.
b. The company introduced a compute multiplier, defined as open platform and API revenue per unit of compute invested across training and inference. This metric improved 14x vs. 1H last year.
4. Narrower losses; management said GP has begun to fund R&D
a. 1H R&D spend was RMB 2.13 bn, approximately $317 mn. Investment remained disciplined.
b. Period loss was RMB 2.072 bn, down 12.5% YoY, with the loss ratio narrowing 4.7x YoY. Adj. net loss was RMB 1.964 bn, with the adj. loss ratio narrowing 3.5x YoY.
c. Management argued that adj. net loss of RMB 1.964 bn being below R&D of RMB 2.13 bn implies GP now covers G&A and sales and has started to fund R&D. This marks a key inflection.
II. Earnings call details
2.1 Management remarks
1. Capability ladder and the commercial formula
a. Management framed LLM evolution as a five-step ladder that must be climbed in sequence: Chat, Coding, Agent, Co-worker, Autonomous AI. Each step has technical gates, and failing to pass means the commercial model at the next step will not be mature.
b. During the period, the focus was between step two and step four, namely Coding through Co-worker. Trials are underway in cyber security, legal, financial, and education, with cyber security advancing fastest while others remain early without scaled revenue.
c. The formula stated: AGI commercial value equals the intelligence ceiling times token consumption scale. Each higher task boundary expands the addressable market by an order of magnitude.
2. GLM-5.3 and GLM-5.3 Flash
a. GLM-5.3 and 5.2 share the same architecture, total parameters, and active parameters, with the only variable being post-training scale. End-to-end task completion improved by over 50% as training shifted from exercise-like tasks to full professional tasks closer to real expert work.
b. GLM-5.3 Flash targets high-frequency, large-scale, cost-sensitive use cases, featuring 320B total parameters and 18B active parameters with linear and hybrid attention. It is priced at one-tenth of 5.2 and outperforms 5.2 in benchmarks and real use.
c. Flash launched anonymously and immediately topped OpenRouter while setting a token usage record at the time. It lifted overall platform call volume by over 20%.
3. Three-stage evolution of the biz. model
a. Pre-2025 focused on on-prem deployment, with customers buying tools to be integrated, customized, and delivered into their own environments. This was coupled with data security compliance and autonomy requirements.
b. In early 2025 the company entered the Coding stage, shifting models from knowledge-driven to task-driven. The firm strategically wound down on-prem software licensing and pivoted to callable, extensible, and metered intelligent services.
c. This year moved into a parallel Agent and Co-worker stage, with the biz. shifting from selling calls to selling subscriptions and end-to-end task outcomes. Management summarized that each capability leap moves purchases closer to ultimate economic value, rewriting the revenue mix each time.
4. Platform users and customer quality
a. As of Aug 2026, MaaS platform registered users exceeded 7.4 mn, up 144% YTD. Paying DAUs rose 603% YTD, and the top-10 customers by revenue saw average daily calls up 98x vs. the start of the year.
b. In just the past two months, driven by GLM-5.2, 5.3, and 5.3 Flash, users rose by 1.6 mn from 5.8 mn at end-Jun to 7.4 mn. Momentum remained strong.
c. By ARR cohorts, customers contributing over $100k, $500k, $1 mn, $10 mn, $25 mn, and $250 mn number 115, 25, 37, 8, 2, and 2, respectively. The mix is broadening upmarket.
5. Safety and governance
a. The approach is to grow open source, openness, and safety in parallel. By open-sourcing weights, the MaaS platform, coding plan, and a global developer network, models enter as many real environments as possible while reliability is built via access control, process supervision, risk assessment, vulnerability disclosure, and external validation.
b. The most sensitive capabilities are controlled, with third parties independently assessing vulnerabilities pre open source and responsible disclosure via the national vulnerability database. The division of labor is that the company self-reports benchmark scores while external teams judge risk, and this does not change with capability gains.
2.2 Q&A
Q: How long can the company sustain its lead in coding, and what are the key variables?
A: Point-in-time leaderboard leads will be matched; the goal is sustained leadership capability. Management noted coding benchmarks update weekly or monthly and are quickly matched, and the leaderboards contain noise. Hence the emphasis is not on how long a SOTA lead lasts but on how strong sustained leadership capability is.
The evidence is continuous multi-gen iteration over roughly the past 11 months from 4.6 and 4.7 to 5.1, 5.2, and 5.3, with the capability curve trending up. Unit task cost has stayed at a relatively healthy level.
Management added that coding is not viewed as a fast revenue growth product line but as a must-clear technical bottleneck. Advancing beyond coding is required to evolve into Agent, long-horizon tasks, and Co-worker.
Q: How will these accumulated capabilities migrate beyond coding, and what is the path to deployment?
A: Migration has been validated in cyber security, with CyberGym rising from 77.2 to 84.5. Coding was chosen because it offers a scalable, automated, low-cost, and verifiable environment where code execution, test pass rates, and bug fixes can be judged objectively. This is critical for RL, while most knowledge work is hard to verify.
Another benchmark improved from 24.4 on 5.2 by 30pts to 54.4, with larger gains in longer, more complete planning tasks. The longer the horizon, the bigger the relative lift for 5.3 vs. 5.2.
In real-world settings, working with multiple domestic security teams, the company identified 2,436 vulnerabilities in real codebases, de-duplicated and screened by experts. Over 1,000 are high risk, spanning 269 projects.
Scenario selection follows three criteria: sufficient intellectual barrier, effective verifiability, and adequate economic value upon completion. Legal, financial, and data analytics are being explored, but reliability thresholds and verification differ, so commercialization structures will differ too.
Q: What is the progress on compute expansion?
A: No card counts; management prefers effective compute and the compute multiplier. Compute supply is an industry-wide issue, and the company is expanding steadily this year with increased investments on both training and inference.
They avoid card counts because sources and uses are diverse, including owned clusters, leases, and procured services. Uses span pre-training, mid-training, post-training, and inference, and efficiency varies widely across chip generations and architectures, so sums do not reflect real capability.
Internally they track installed compute that can be stably scheduled and run long-term, and ultimately converted into training progress and billable tokens. The real question is not the number of cards but the model iteration speed and the scale of commercial token supply those resources support.
Q: How much model training and inference can current compute support?
A: Inference has reached a 100k-card class of domestic chips; training remains a heterogeneous mismatch. A 100k-class domestic chip cluster now supports scaled, low-cost inference, cutting unit inference cost by 80% vs. early-year.
When GLM-5.3 Flash launched anonymously, all services ran on domestic clusters, with about 60 trillion tokens served in six days, a record for both OpenRouter and OpenCode. Management believes domestic cards can run the workloads, and the deeper test is economics.
To that end, the inference engine and service stack were rebuilt over the past half-year, including quantization, PD decoupling, layer split, caching, and communication optimizations. End-to-end performance improved 3x vs. baseline on the same domestic hardware.
On training, the issue is structural rather than quantity, as homogenous clusters demand high stability, networking, and software stack maturity, while domestic capacity is mostly heterogeneous. Management expects improvement in 3–6 months as advanced domestic cards scale, and base model training is underway with compute prioritized for critical training windows rather than diverted to inference ramp.
Q: Will models commoditize with value shifting to the harness or token orchestration layer?
A: They agree point leads do not confer pricing power; four levers sustain advantage, and two price curves will emerge. The four are iteration speed, real task completion rate, unit intelligence cost, and workflow depth with customers.
On speed, six generations in 11 months lifted the intelligence index from 32 to 60, with flagship single-task costs around $0.2 and Flash around $0.045. Point scores can be copied, but this curve is harder to replicate.
On completion, converging public configs will bring single-question ability closer, but longer tasks widen gaps. Multi-hour engineering tasks require reading codebases, tooling, failure handling, replanning, and delivery, and value density differs widely between 1 mn tokens spent on chit-chat vs. production bug fixes.
On workflow depth, customers build prompts, permission systems, eval standards, and engineering integrations around the model, raising switching costs. Evidence shows avg. API prices rose about 101% while volume still surged, with top-10 customer daily calls up 98x vs. the start of the year.
Thus, they expect prices for same-tier intelligence to keep falling while models that expand task boundaries retain premiums, even room for price hikes. Dual curves will likely coexist.
Q: What are the developments in overseas competition and go-to-market?
A: Overseas has shifted to the open platform and API model; CSP partnerships are expected within 1–2 months. Since 2025 the company has had overseas business of meaningful scale, initially focused on on-prem sovereign LLM projects such as the previously disclosed Malaysia project.
With the pivot to open platform and APIs, GLM-5.3 Flash is a first attempt under the new model. Management acknowledged that Chinese models still trail top overseas peers at the frontier.
However, on a combined intelligence and cost basis, GLM-5.3 represents high-end intelligence while 5.3 Plus represents broad, high value-for-money, occupying two ecological niches. The dual Pareto chart in the interim report illustrates this.
They will push collaborations with overseas CSPs, and may consider local hosting of open-source models and revenue-sharing arrangements going forward. More formats will be explored.
Q: What are the core improvements for the next-gen model, and how will parameters, architecture, and training change?
A: Scale the base further, achieve deep reasoning with small activation, and strengthen post-training, with a goal of fully self training. The technical path has converged, with earlier investments in multimodal, video, and image generation found to add little to the intelligence ceiling, so the focus is on text.
The base will continue scaling to store and better represent knowledge, raising the intelligence ceiling. But activation cannot simply be enlarged, as too large activation slows inference and raises cost, so the direction is small activation with new methods to achieve deeper reasoning.
Post-training is a self-assessed strength. China cannot match the compute of top US labs in the near term, so new technical and engineering methods are necessary, and since Nov 2022 the company has built a data labeling team of several hundred, with 5.3 gains mainly from a many-fold expansion of the data environment.
The longer-term target is fully self training, also known overseas as RSI, recursive self improvement, where the model trains itself from pre-training through mid- and post-training. The biggest challenge is whether the model can decide when to stop and when to self-correct, and ethics and social governance will be embedded in the next generation.
Q: Some peers are pushing to even larger parameter counts. Why is the company more restrained, and what is the plan?
A: 5.3 base has 745 bn parameters with a design life of over half a year; next-gen parameters are undisclosed. Parameter choice balances three factors: available resources, the desired end-state of the model, and when users can access it.
On data, current domestic datasets are generally 30–50T, and with both data and compute constrained, the optimal size is debatable. Management also noted they have not pushed existing resource limits to the max, so piling on parameters is not the most critical task.
Since the start of the year the roadmap has been clear: in Jan they set 5.3 base at 745 bn parameters with a plan to use it for over half a year, defining what 5.1, 5.2, and 5.3 must achieve and the corresponding data environment scale and campus planning.
The next gen will also have a very large base, but parameters will not be disclosed, and the specific mid- and post-training on top are under discussion. The principle is to amplify each base to the maximum while making each generation truly usable, and inference cost has been optimized so enterprise open-source deployments can run practically, with next-gen inference optimization already underway.
<End of text>
Risk disclosure and disclaimer:Dolphin Research Disclaimer and General Disclosure
Login to unlock14,635characters for free
This content is only available to signed-in users. Sign in to your Longbridge account to read the full post.
