Databricks has released an AI entity named KARL. The core goal is to reduce ineffective searches in retrieval-enhanced generation systems. The company stated that such systems used to continuously pull context and repeatedly retrieve information until they reached the token limit or timed out, which not only increased computational power consumption but also slowed down response times.
Emphasize when to stop searching
The design focus of KARL is not merely to expand the search scope, but to determine when sufficient information has been obtained. Databricks indicates that this capability is trained through reinforcement learning, enabling the agent to stop actively when further searches can no longer improve the quality of the answer, rather than continuing to consume resources.
The company also combines "context compression" with this mechanism. In other words, KARL will compress the information that has already been obtained before deciding whether to continue searching. Databricks believes that these two capabilities together form the basis for KARL to improve efficiency.
Cost and latency data release

According to Databricks, KARL can achieve an accuracy rate on retrieval and inference tasks that is on par with Claude Opus's 4.6, but it has a lower operating cost and faster response times.
- Cost reduced by 33%
- Delay reduced by 47%
- The accuracy rate is comparable to Claude Opus's 4.6.
This means that when dealing with a large number of enterprise query scenarios, both system deployment costs and waiting times can be further reduced. For enterprises that need to operate AI intelligents on a large scale, such indicators are usually of more concern than a single benchmark score.
Connect to the Agent Bricks platform
KARL is not an independent product, but rather runs within the Agent Bricks platform launched by Databricks in September 2026. This platform is designed to build automatically optimized, scenario-specific agents for enterprises, with the goal of reducing the workload involved in manual parameter tuning for companies.
Databricks indicates that Agent Bricks supports mainstream models such as Claude and GPT, and processes governance through Unity Catalog. The company has also added reranking functionality to AI Search products, which involves re-scoring preliminary search results before passing them on to the language models for processing. Databricks claims that this feature has increased the accuracy of enterprise benchmark tests by about 15 percentage points.
The platform has over 100,000 agents in use.
According to Databricks, since the launch of Agent Bricks, over 100,000 agents have been built on the platform. The company also stated that by combining methods such as parallel reasoning and multi-model collaboration, the accuracy for certain types of tasks has increased from around 32% to over 90%.
From a product positioning perspective, Databricks does not define itself as a basic model provider, but rather emphasizes providing the infrastructure necessary for enterprises to deploy, govern, and scale models. The release of KARL also continues this approach: the focus of competition is not just on the models themselves, but on their operational efficiency and usage costs within the corporate environment.









