On September 15th, Google updated AI and Economy ATLAS, and announced a study in collaboration with Google DeepMind and MIT FutureTech. The study analyzed 2,600 professional AI models and surveyed over 600 scientists from the United States and the United Kingdom. The results showed that nearly half of the surveyed scientists used some form of AI daily, and they reported saving nearly 7 hours per week as a result. However, the study also found that with the rapid generation of more hypotheses, there was a backlog in physical experiments, clinical validation, and result verification. This conclusion is more comprehensive than the claim that "AI can speed up scientific research": while models can streamline some mental processes, they cannot automatically expand laboratory, sample, and approval resources.
Seven hours is not an automatically increased amount of scientific research output, but rather time that has been reallocated.
Scientists do not use AI in a single way. Large language models span multiple disciplines and tasks, being used for searching for information, organizing thoughts, writing code, and processing text; specialized models are more focused on life sciences, health research, as well as specific data prediction, generation, and simulation. These two types of tools complement each other: general-purpose models lower the barriers to expression and programming, while specialized models delve deeper into domain knowledge and data structures. To categorize all uses of AI as "chatbots" would overlook the dedicated systems that truly impact the scientific research process.
The nearly 7 hours reported by the respondents are self-reported savings and do not equate to the average productivity verified through random trials, nor do they represent that all disciplines have achieved the same benefits. Teams familiar with programming and having ample data may be able to integrate AI into their workflows more quickly; however, in fields that rely on expensive equipment, require long-term follow-ups, or involve on-site sampling, the time saved at the front end may not necessarily shorten the overall project duration. The survey participants mainly came from the United States and the United Kingdom, so the results may not directly represent the global scientific community.
The most inspiring part of the research is the “hypothesis backlog.” As literature summarization, data analysis, and candidate generation speed up, researchers can propose more questions worth testing. However, experimental equipment, ethical approvals, clinical samples, and the judgment of senior researchers do not expand in tandem, so the bottleneck shifts from idea generation to verification. A common pattern seen in enterprise software also occurs in laboratories: optimizing one step often merely shifts the queue to the next step.
This means that evaluating the AI research tool cannot be limited to calculating how many minutes are saved in writing or coding. More importantly, it is about verifying how many valuable hypotheses were ultimately confirmed, whether the number of incorrect candidates decreased, whether the failure rate of experiments reduced, and whether the results of papers and data can be reproduced. If a model generates ten times as many candidates for the team, but there is no better mechanism for sorting and eliminating them, the experimental burden may actually increase.
Google also launches a new ATLAS interactive experience, allowing users to see the usage of AI in different occupations, countries, and tasks. For example, in OECD countries, professions such as computing and mathematics, as well as business and finance, lead in usage; in non-OECD countries, occupations like office administration, creative media, and education are more prominent. In the United States, among work-related AI usage, computing and mathematics professions account for 30%, which is approximately twice the proportion in other regions. These data describe the composition of usage and should not be misinterpreted as meaning that 30% of all American technicians are using AI.
What scientific research institutions need to transform is the verification system, not just creating an account for each individual.
The first step is to record AI separately from the research evidence. Literature summaries, code, data cleaning, and candidate predictions should retain the model version, hints, source of input, and manual modifications. Research conclusions must be based on the original data, experiments, and peer review. Tools can provide explanations, but a high confidence level should not be regarded as a level of evidence.
The second step is to invest in downstream capabilities. If generation speed is assumed to be faster, institutions should increase resources for shared experimental platforms, automated instruments, data engineering, and statistical review, and establish priorities for candidates. Otherwise, the seven hours saved will turn into more waiting time. Managers need to observe the overall changes in project cycles, rather than just looking at the number of times employees use these resources.
The third step is to identify biases. Professional models rely on existing data, and rare diseases, low-resource languages, and populations with insufficient representation are more likely to be overlooked. Large models may also fabricate citations or repeat existing mainstream assumptions. Teams should maintain independent search, negative result, and counter-evidence processes to prevent efficiency tools from further magnifying the blind spots in academic consensus.
The fourth step is to adjust talent training. If young researchers directly obtain the organized answers, they may skip the process of understanding the experimental design and data limitations. A reasonable approach is not to disable AI, but rather to require them to explain their choices, review citations, reproduce experiments, and invest the saved time in more in-depth learning in the field. Only those who can identify model errors can use the models safely.
ATLAS itself is a long-term research project, and Google indicates that efforts will continue to expand with academic partners. The current results provide information on usage patterns and respondents' perceptions, rather than representing a certain proportion of scientific discoveries that AI has already led to. In particular, aspects such as "daily use" and "time savings" may be influenced by sample selection, task definition, and self-reporting, and therefore require independent verification through subsequent research.
This research takes the discussion on AI productivity one step further. The real question is not whether scientists can write code or summarize papers faster, but whether the entire discovery chain can accommodate this increased pace. When experiments, clinical trials, and reviews cannot keep up, what models produce is a pool of candidates, not knowledge. In the future, the most successful scientific institutions may not be those that generate the most hypotheses, but rather those that are best at screening, verifying, and making uncertainties public.










