Clockwork.io Raises $31 Million in Financing
PR Newswire
1h ago
Ai Focus
Clockwork.io Announces $31 Million in Financing; LinkedIn, Together AI, and WhiteFiber Are Adopting Its Fault-Tolerant Software to Reduce GPU Idle Time. The Company Also Released New Features for TorchPass to Preserve the Status of Distributed AI Tasks Without Changing Code, and to Accelerate Reinforcement Learning.
Helpful
No.Help

Clockwork.io Raises $31 Million in Financing

– Clockwork.io raised $31 million in financing; LinkedIn, Together AI, and WhiteFiber adopted its fault-tolerant software to avoid GPU hours of wasted time.

LinkedIn can prevent tens of thousands of hours of GPU downtime each month; the new innovations of TorchPass allow us to maintain the AI workload progress without changing the code, while also accelerating reinforcement learning.

Palo Alto, California, USA, October 5, 2026 / PRNewswire / — Fault-tolerant software provider Clockwork.io, which ensures that AI training, reinforcement learning, and inference workloads continue to operate in the event of infrastructure failures, today announced that it has received $31 million in new financing. The company has also completed production deployments in LinkedIn and Together AI, while further expanding its adoption in WhiteFiber.

The company has also released two new features of the TorchPass solution, each of which can capture the status of running distributed AI tasks. The multi-mode platform snapshot is an industry-first in the field of training; it allows for the preservation of complete tasks on all nodes without the need to modify the training code, and these tasks can be restored later. The fast asynchronous application checkpoint feature is executed in the background while the tasks are running, which can accelerate reinforcement learning. These features provide the updated model weights to the inference copies responsible for generating rollouts, which are the samples upon which the model learning is based. As a result, these inference copies have reduced waiting times or avoid working with outdated models.

Fault tolerance has become a basic requirement for large-scale AI.

Large distributed AI workloads can span thousands of GPU. These GPU must be kept synchronized: a failure of a single GPU, a disruption in the link, or a server crash can all cause the entire task to come to a halt. Meta reported that during a 54-day training period using 16,384 GPU, unexpected interruptions occurred on average about every three hours.

The usual approach is to reload the checkpoint, which is a saved copy of the task progress. The recovery process can take up to 90 minutes, causing the normally running GPU to be idle, and forcing the system to repeat the work that has already been completed from that checkpoint. Customers have to bear the cost of the idle GPU and the additional calculations, and the model completion time will also be longer as a result. As the task scale expands, each restart puts more GPUs at risk of increased usage time.

At this scale, it has become a requirement of infrastructure to ensure that work continues to progress even in the event of a failure. Clockwork.io meets this need through a set of fault-tolerant tools, which the platform team deploys as an intermediate layer between hardware and workloads. LinkPass re-routes traffic to avoid faulty links, so that tasks are not even aware of the failure. TorchPass allows workloads to be migrated from faulty GPU to normal GPU, enabling training to continue without having to revert to an earlier state. Both solutions have now entered the production phase. The new TorchPass platform snapshot feature announced by the company today can capture the status of running distributed tasks; when a failure is severe enough to cannot be resolved by simple migration, this feature can be used to restore the entire task.

"At the scale of AI, failures are inevitable; however, losses of several hours of productive work should not occur as a result," stated Suresh Vasudevan, the CEO of Clockwork.io. "Fault tolerance is like an effective performance amplifier: it allows GPU to continue working productively, rather than waiting for recovery processes or repeating tasks that have already been completed. We collaborate with companies that operate some of the largest GPU clusters and cloud service providers to develop software, so it can handle the failures they actually encounter. This kind of protection should become part of the infrastructure that these companies and service providers rely on daily."

The scope of adoption has been extended to enterprises, ultra-large-scale clouds, and neocloud.

Enterprises that build their own GPU clusters, ultra-large-scale cloud vendors, and “neocloud” providers adopt Clockwork.io for the same reason: to make more of those GPU hours available for effective work.

LinkedIn has deployed LinkPass network fault tolerance capabilities across its entire infrastructure based on AI, thereby avoiding tens of thousands of hours of GPU downtime each month.

LinkedIn, Senior Vice President of Infrastructure and Chief Technology Officer, Raghu Hiremagalur stated: "With the infrastructure scale of AI, a single network issue should never cause a normally functioning GPU to go offline or interrupt ongoing workloads. Before adopting Clockwork.io, fluctuations in InfiniBand NIC could cause a server equipped with 8 GPU to fail, and fluctuations in switch ports could cause a second server to also fail, doubling the impact to 16 GPU. Clockwork.io has helped change this operational model. Its network fault-tolerance technology automatically redirects traffic to available paths, allowing tasks to continue running during repairs to links, optical components, cables, or network cards. Overall, Clockwork.io has prevented tens of thousands of hours of GPU downtime for our entire cluster each month. It has transformed what were once destructive operational events into manageable maintenance events, helping to improve infrastructure utilization and operational efficiency."

Together AI is launching TorchPass as a service on its GPU cluster. During the PyTorch conference, both parties will demonstrate a multi-node training task on-site; even if network and GPU failures are intentionally introduced, the task will continue to run without the need for a restart.

"The standard by which customers measure us is goodput, which is the proportion of time within GPU that truly contributes to the progress of the model," said Together AI, the product leader for Pavneet Ahluwalia. "The node repair function is now capable of automatically detecting failures and allocating alternative capacity. TorchPass and LinkPass are built on this infrastructure, aiming to ensure that tasks can continue to execute even in the event of GPU and link failures, thereby preserving the progress that has been made. We are now bringing these features to market as the next level of fault tolerance for our platform."

WhiteFiber ( NASDAQ : WYFI ) is an existing customer of the company and is expanding the use of Clockwork.io software within its continuously growing global GPU service infrastructure.

WhiteFiber Chief Technology Officer Tom Sanfilippo stated: "It is crucial to conduct stress tests on the reliability of a cluster before putting it into production, because if customers receive a system with hidden interconnection network failures, they will ultimately face the consequences of failed tasks and lost GPU hours of work. Optical components of average quality, improperly configured network cards (NIC), and links that pass basic tests but degrade under load can all be overlooked. Clockwork.io's automated cluster auditing verifies each link and each node simultaneously, locating faults within minutes and allowing us to fix them before delivery. We can launch clusters more quickly, so customers' first training runs are conducted on a network that has undergone end-to-end verification, rather than just on a system that is already up and running. Given the rapid market demand, delivering verified capacity to customers promptly is vital to our business; therefore, we are deploying Clockwork.io in all of our clusters."

The new features of TorchPass hand over workload protection to the platform team.

Clockwork.io has extended the capabilities of TorchPass from GPU to include two functions that the platform team did not previously possess: the ability to generate snapshots of entire distributed tasks on their own, as well as fast, background-operable application checkpoints.

In training scenarios, platform snapshots save the execution status of running tasks on all nodes, allowing for task recovery after interruptions. The platform team and AI infrastructure engineers can deploy this solution on compatible workloads without requiring application owners to modify code or add checkpoint logic. Enterprise teams can protect training tasks across the entire fleet through a single mechanism, and cloud service providers can also safeguard customer tasks for which they do not have control over the code. When teams implement checkpoints at the application layer, TorchPass's fast checkpoint capability enables more frequent checkpoints; as a result, less progress is lost in the event of a failure, and less computational work needs to be redone.

In terms of reasoning, large models typically run on two or more servers; therefore, a link failure can render the entire replica inoperable and interrupt user sessions or proxy tasks mid-way. LinkPass allows these multi-server replicas to continue running in the event of a link failure.

Reinforcement learning relies on both training and inference. Multiple model copies generate rollouts; the training system learns from them, and the updated weights must first be applied back to these copies before they can use the latest version to generate results. The application checkpoints of TorchPass allow for faster transmission of updated weights back to the executing copies, while LinkPass ensures that these copies continue to operate in the event of link failures.

SemiAnalysis Founder, CEO, and Chief Analyst Dylan Patel stated: "Cluster fault tolerance used to be just a training issue; now it has also become a problem for inference. According to our tests with ClusterMAX, at a neocloud provider that has received a 'gold medal' rating, TorchPass reduced the effective performance loss during training from 14% to less than 3%. Reinforcement learning (RL) connects these two processes: inference replicas generate simulations, the training system learns from them, and the updated weights are then fed back to the replicas. Clockwork.io enables the replicas to continue running even in the face of link fluctuations and network failures. Its ultra-fast checkpoints also accelerate the transmission of weights back to the inference cluster, preventing the process from stalling in either direction. This solution needs to be deployed as a general fault-tolerant layer for training, inference, and reinforcement learning."

Enterprise platform teams and cloud service providers can contact Clockwork.io to evaluate the application of their software in their respective workloads or to discuss potential collaboration opportunities.

Financing Round and Use of Funds

This round of financing was co-led by Premji Invest, Wing Venture Capital, and Seligman Ventures. Existing investors NEA and e& Capital also participated. With this, Clockwork.io has raised a total of 73 million US dollars. The company will use these funds to accelerate the deployment of its fault-tolerant tool suite, which is suitable for training, inference, and reinforcement learning phases, to increase corporate adoption, and to expand delivery scale through cloud partners.

"The only thing that can be perfectly scaled is unreliability: as long as you install enough GPU on one machine, there will always be problems," said Greg Papadopoulos, a venture capital partner at NEA. "Traditional methods—stopping tasks and reloading checkpoints—are no longer meaningful at today's scale. Clockwork treats failures as the norm: TorchPass allows training to be seamlessly migrated from faulty GPU systems, and it is now possible to capture the state of an entire running task without having to modify the code. It has already saved tens of thousands of hours of GPU time each month. We invested in Clockwork.io in 2021 and are delighted to continue supporting them in defining the performance layer of their AI clusters."

About Clockwork.io

Clockwork.io is the pioneer of Software - Driven AI Fabrics ™. It represents a programmable layer situated between hardware and workloads, enabling GPU clusters to achieve observability, fault tolerance, and maximize utilization, regardless of the impact of the accelerators, networks, or cloud environment in use. AI workloads require the entire cluster to function as a single machine; however, failures and bottlenecks can cause GPU to become idle. The FleetLens platform of Clockwork has the capability to recover from such losses: with nanosecond-level telemetry, it can accurately identify which GPU, which node, or which link is slowing down or blocking tasks, and it can verify the cluster before it goes live. Furthermore, through LinkPass and TorchPass, it ensures that workloads from training to inference can continue to run even in the event of infrastructure failures or degradation. Independent benchmark tests conducted according to SemiAnalysis have shown that TorchPass outperforms major open-source frameworks as well as traditional checkpoint and restart methods in terms of fault recovery speed. Companies such as LinkedIn, Together AI, WhiteFiber, Wells Fargo, Nebius, NScale, and DCAI are all using Clockwork.io to drive their AI infrastructure. For more information, please visit www.clockwork.io.

Tip
$0
Like
0
Save
0
Views 19
HQYC reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Citi is optimistic about the continuation of the global stock market rally, stating that the S&P 500 and the Dow Jones Industrial Average may still continue to rise by 2027
Citi strategists stated that since 2026, global stock markets have risen by about 12% and are still near historical highs. Although the MSCI global index is expected to achieve double-digit gains for the fourth consecutive year, Citi believes that the current market is more characterized by "resilience" rather than "complacency." They also pointed out that corporate earnings remain the core factor supporting stock markets, but rising U.S. Treasury yields remain a major risk.
Coinpaper
·2026-10-06 01:11:03
5
MN ETF seeks to enable investors to access OpenAI and Anthropic
Corgi Invest Announces the Launch of MN (Corgi MANGOS ETF), an actively managed exchange-traded fund aimed at enabling investors to access OpenAI and Anthropic, as well as Meta Platforms, NVIDIA, Alphabet (Google), and SpaceX through methods such as cash-settled total return swaps. The fund began trading on Cboe BZX exchange on October 2nd, with an annual total operating expense rate of 0.20%.
PR Newswire
·2026-10-06 01:10:59
4
FF EAI Robotics Ecosystem Inc Claims Record-Breaking Robot Sales and Shipments in September, Reaching 265 Units
FF EAI Robotics Ecosystem Inc stated that in September, the sales and shipments of its EAI robot devices reached 265 units, setting a new monthly high; as of the third quarter, the cumulative total was 575 units, and from the end of February to the end of September, it totaled 817 units. The company also mentioned that on October 14th, they will jointly host a new product launch event with FFAI, and continue to advance the independent listing of their robotics business.
PR Newswire
·2026-10-06 00:40:05
13
Tom Lee Says the Rise in the Crypto Market Signals a Shift: The Fed Will Move Towards Loose Monetary Policy, and Democrats Will Support the CLARITY Bill After Midterm Elections
The Tom Lee of Fundstrat stated in an interview with CNBC that the strength of crypto assets and tech stocks indicates that the market does not believe that a tighter monetary environment is on the horizon. He expects that as inflation subsides and financial conditions improve, the Federal Reserve will move away from its hawkish stance; at the same time, Democrats will shift to support the CLARITY legislation after the mid-term elections.
Coinpedia
·2026-10-06 00:00:17
28
View More