LegalOn Reduces Development Costs with Three Models: The Key to Saving Money Isn't "Replacing All with Smaller Models"
CoinMeta
1h ago
Ai Focus
After enterprises implement the AI programming assistant within their R&D teams, the billing process typically becomes much more complex than during the demonstration phase. The same development task can result in varying usage of the model depending on factors such as the clarity of requirements, the size of the codebase, and whether architectural decisions are necessary. In a case study published on October 8th by OpenAI, this legal technology company demonstrates how it manages Codex costs while maintaining development speed: they select models based on the complexity of the tasks, limit the default use of the fast mode, and set different budgets for various business areas. According to the case study, the estimated daily cost has decreased by about 65% compared to the previous method of usage. This figure is quite impressive, but it cannot be taken out of context...
Helpful
No.Help

After enterprises implement the AI programming assistant within their R&D teams, the billing process typically becomes much more complex than during the demonstration phase. The same development task can result in varying usage of the model depending on factors such as the clarity of requirements, the size of the codebase, and whether architectural decisions are necessary. In a case study published on October 8th by OpenAI, this legal technology company demonstrates how it manages Codex costs while maintaining development speed: it selects models based on the complexity of the tasks, limits the default use of the faster mode, and sets different budgets for various business areas. According to the case study, the estimated daily cost has decreased by about 65% compared to the previous approach. This figure is impressive, but it cannot be sold separately without considering the underlying calculation methodology.

LegalOn initially allowed developers to use high-performance models more freely, integrating these tools into design, implementation, and daily work. As the scale expanded, the question was no longer whether AI was useful or not, but rather whether limited budgets were being spent on the areas that most urgently needed strong models. Internally, the AI Center of Excellence for Development tested and monitored different models, then provided selection guidelines to each team. With clear requirements, Luna was used more for implementation, Sol for routine design, analysis, and documentation tasks, while Astra was reserved for complex architectures and advanced decision-making. The company also placed monthly usage limits on departments and individuals under the control of administrators. It is these combined measures that constitute the context of cost variations.

Layering tasks is more difficult but also more effective than a one-size-fits-all downgrade approach.

Model routing may sound simple, but in reality, it's necessary to first define the boundaries of the tasks. Writing an interface with established specifications is one responsibility, while deciding how the system should be divided, how to handle permissions, and how to manage failures are other responsibilities. For the former, low-cost models can be tried out, followed by testing and review; however, if the wrong tools are chosen in an attempt to save on call costs, the cost of rework could far exceed the difference in model prices. What can be learned from the LegalOn case is not some permanently valid list of models, but rather knowing when to upgrade to a lightweight model and when it is necessary to seek human judgment for final decisions.

Another change that is easily overlooked is the fast mode. The case study mentions that companies typically limit its use, allowing it only for individual requests when truly necessary. Accelerated generation can be valuable in emergency troubleshooting and exploration, but treating every daily task as a high-priority one is like sending all emails through express delivery. Managers need to consider both the waiting time and the total completion time: if the normal mode is slightly slower but does not delay delivery, the savings in cost are substantial; however, if the restrictions cause developers to frequently queue or submit requests repeatedly, the decrease in costs may mask losses in productivity. LegalOn states that the team maintained speed by working on tasks in parallel, which is still a valuable experience for that company, but external teams should measure these effects for themselves.

The allocation of resources within a company is not evenly distributed among individuals. As mentioned in the OpenAI case, mature businesses are required to improve their cost efficiency by about 20%, while new businesses in the startup phase are given more flexible budgets. There is a logical reason behind this distinction: established products already have a baseline of revenue and costs, allowing managers to compare input with output; new products are still seeking a market, and it would be inappropriate to impose a uniform budget limit too early, which could stifle the space for experimentation. If the “departmental budget limits” are simply applied without considering the stage of the business, it is possible that teams that need experiments the most will be the first to face restrictions.

65% of these numbers still need to be read separately. The source states that this represents a reduction in daily costs compared to the previous usage method of GPT-5.5, which is due to a combination of model selection, quick mode adjustments, and budget arrangements; it is not a guarantee that all companies will see the same reduction in cost per unit of code, nor does it necessarily mean that the cost will decrease by 65%. There is also about a 20% improvement in efficiency in mature businesses, which is measured differently from the overall estimated daily cost changes. Confusing these two figures can create the illusion that "turning off one feature can save a large portion of the bill." Financial officers should ensure that the task volume, delivery volume, failure retry rates, and manual review times for the same period are all accounted for in the financial records.

The real metrics that should be tracked are function delivery, not the number of calls.

In the LegalOn case, the most noteworthy remark is actually the reservation: the company believes that faster development does not automatically equate to higher customer value. This is very realistic. A feature can be developed in two days, but if users do not need it, if the quality is unstable, or if the maintenance costs increase, then the savings in model fees and man-hours cannot translate into better business results. When measuring the programming investment in AI, one must observe at least the usage of the launched features, defects, rollbacks, developers' waiting times, and the total costs. Focusing solely on the number of lines of code generated or the number of interactions can easily lead to rewarding ineffective work.

If a R&D supervisor wants to verify similar strategies, they can first select two to three types of high-frequency tasks, record the original time and cost involved, and then compare different model combinations under the same acceptance criteria. Fixes with clear requirements, cross-module designs, and fault investigations should be separated; one average value cannot represent all of them. The test set should also include failure cases: whether the model accidentally modified unrelated code, whether security boundaries were ignored, and how long it took for engineers to correct these issues. Only when both the savings in cost and the reduction in rework are significant can the strategy be considered effective. If a cheaper model requires senior engineers to spend more time to complete it, the low price on paper is meaningless.

This is still a case study published by a supplier, based on the practices of a single client, and not an independent experiment across industries. It illustrates how LegalOn organized model selection and budgeting, and provides its own estimated results; however, it cannot be inferred that the legal technology industry will all achieve similar benefits. At the same time, the models and prices mentioned in the case study may change with product updates, and teams should not regard the configuration from October 2026 as a long-term fixed standard. What can be replicated are the classification, measurement, and feedback mechanisms, rather than the specific model names at any given time.

For companies that are developing tools to scale up for AI, this case provides a practical sequence: first, clarify the tasks and risks; then, match the model capabilities with the costs; finally, assess the customer outcomes after release. The most expensive models are not always necessary, and the cheapest models should not become the default management goal either. Organizations that truly reduce costs often do so not by decreasing the number of calls, but by spending their limited budgets less on areas where they are not needed.

Tip
$0
Like
0
Save
0
Views 21
HQYC reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Under pressure of performance, a large purchase of 20,000 NVIDIA GPU Tianyang Technology shares led to a price increase of over 12%
On October 9, Tianyang Technology disclosed that it had signed a GPU procurement agreement with Company X to purchase 20,000 NVIDIA RTX PRO units at 5,500 Blackwell each, resulting in a 12.38% increase in the closing stock price. The article stated that this product is not a AI training chip for data centers, and based on market prices, the value of this procurement is estimated to be between 1.6 billion and 2.4 billion yuan.
The Block
·2026-10-09 16:42:57
7
What will happen to the stablecoins still held by customers if Europe demands their withdrawal from exchanges?
ESMA has required regulatory authorities around the world to resolve the existing positions of stablecoins that do not comply with MiCA regulations within three months. For regulated crypto companies, related services must be ceased, but under strict supervision, existing positions can still be liquidated, converted, withdrawn, or transferred.
crypto.news
·2026-10-09 16:42:55
4
XRP Testing the $1.32 support level, ETF Demand slowing down
After failing to break through the $1.60 resistance level, XRP has seen a cumulative decline of nearly 7% in 7 days. Analysts note that since October, XRP ETF has only recorded approximately $4 million in net inflows, which is significantly lower than the $121.4 million in September; meanwhile, the inflows into exchanges for XRP have risen to their highest level since July 2026, increasing concerns in the market about profit-taking.
CoinJournal
·2026-10-09 16:35:48
9
CaixaBank Expands Strategic Cooperation with Google Cloud to Accelerate the Construction of AI and Data Capabilities
CaixaBank Announces Expansion of Strategic Agreement with Google Cloud, Extending Cooperation Until 2033, Focusing on Advancing Artificial Intelligence, Data Management, Data Analysis, and Cloud Infrastructure Capacity Building. According to the company, the new agreement will incorporate Gemini Enterprise, and will revolve around three main pillars: data analysis and proxy AI, hybrid infrastructure and security, and employee training and empowerment.
PR Newswire
·2026-10-09 16:12:09
20
Apple reportedly cuts orders for iPhone 18% to 20% of Pro series components; rising prices of memory chips drag down demand for new devices
Insiders say that Apple has informed some suppliers to cut the production of iPhone by 18%, Pro by 18%, and iPhone by 18%. The order volume in October has decreased by 15% to 20% compared to the original plan. Reports indicate that rising prices of components such as storage chips have pushed up the selling prices of new devices, and coupled with demand falling short of expectations, this has prompted Apple to lower its shipment forecasts. Additionally, the competition in the AI industry for storage capacity may continue to put pressure on the costs of consumer electronics until 2027.
The Block
·2026-10-09 15:52:01
22
View More