Writer Unveils Model That Reduces AI Agent Costs by 52%
Writer, an enterprise AI agent platform provider, has announced its new flagship model "Palmyra X6". According to the company, when combined with the improved agent foundation, it can reduce AI agent operational costs by an average of 52%. Palmyra X6 is based on Z.ai's open-weight model "GLM-5.2" with Writer's own post-training applied, and the company has disclosed this fact in a technical report.

Writer, a provider of enterprise AI agent platforms, has announced its new flagship model "Palmyra X6". At the same time, the company has unveiled a revamped "orchestration foundation" that consolidates agent operations and a governance tool that allows IT administrators to manage token consumption. Writer's customers include major enterprises such as Accenture, Uber, and Vanguard.
An AI agent is a system that automatically repeats multiple processes—planning, information gathering, tool invocation, verification, and retry—in response to a single user request. Unlike chatbots, which return a single answer to a single question, agents loop multiple times internally, making token consumption, which is the billing unit, prone to increase. While users receive what appears to be a single response, the billing reflects all underlying processes. This structure has become a major factor putting pressure on enterprise AI budgets.
According to Writer, combining Palmyra X6 with the new agent foundation reduces costs by an average of 52%, improves processing speed by 48%, and improves output quality by 10%. Waseem AlShikh, CTO of the company, stated: "Enterprises want to expand token consumption. That's a sign that AI adoption is growing, but at the same time, costs must be controlled." Additionally, Matan-Paul Shetrit, a product management director, noted that cost issues have become the biggest barrier to expanding AI adoption in enterprises, even more so than model capabilities.
Palmyra X6 is not a model trained from scratch, but rather Z.ai (formerly Zhipu AI), based in Beijing, has released an open-weight mixture-of-experts model called "GLM-5.2", to which Writer has applied its own post-training. Writer has disclosed this fact in a technical report. Dan Bikel, head of AI research, explains: "It is ultimately a Palmyra model, using GLM-5.2's numerical parameters as a starting point and building upon it through our own training." Shetrit also emphasized: "It has no connection whatsoever with the original developers and operates entirely on U.S. domestic infrastructure."
There is divided opinion within the industry about an American company using a Chinese open-source model as the foundation for commercial AI. Writer's transparent disclosure of this choice has some significance from a trust perspective. However, how enterprise customers will perceive this requires observation of future developments.
The rapid increase in token consumption is also an industry-wide challenge. According to Goldman Sachs analysis, token consumption is expected to increase 24-fold between 2026 and 2030, reaching 120 quadrillion tokens per month. The main driver of this increase is not an increase in the number of users but rather the proliferation of always-on enterprise agents. The analysis further notes that even if the per-token price decreases, if agents use more tokens, total costs could rise, suggesting that simple price competition alone will not solve the problem.
What distinguishes Writer's announcement is that it offers not only improved model performance but also cost management and governance as a package. This reflects a situation where the competitive axis of enterprise AI is shifting from "what can it do" to "how much does it cost," and suggests that cost efficiency and governance infrastructure are becoming important factors in adoption decisions.
This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.