BUSINESS Overview We are building the fastest AI infrastructure in the world. In AI, speed is critical to win. Speed improves user engagement, expands product capabilities, can lower operating costs, and opens new markets. It shortens iteration cycles for engineers, researchers, and professionals across industries, allowing them to be more productive. Speed unlocks new applications and new industries. In technology, “speed unlocking value” is a pattern that has repeated itself over the past 30 years. Faster solutions are used more often and for more demanding tasks. For example, the speed of broadband transformed the internet from static pages into real-time applications, enabling new products and industries. Similarly, in search, Google showed that even short delays in delivering answers significantly reduced usage and engagement. AI repeats this pattern. As AI has moved from novelty to necessity, AI work has grown more demanding, and speed has become a bottleneck. Faster AI does more work in less time, providing better answers sooner. Our solutions are built for speed. Cerebras Inference delivers answers up to 15 times faster than leading GPU- based solutions as benchmarked on leading open-source models. Similarly, many customers have achieved more than 10 times faster training time-to-solution compared to leading GPU systems of the same generation. These performance breakthroughs are the result of our core innovation: the world’s first and only commercialized wafer-scale processor. Called the Wafer-Scale Engine (“WSE”), our processor is 58 times larger than NVIDIA’s B200 chip and has 2,625 times more memory bandwidth than NVIDIA’s B200 package, which contains two individual chips. To build the WSE, we solved the 75-year-old compute industry problem of wafer- scale integration to produce, yield, power, and cool a chip of this size. This size is what enables our incredible AI speeds. By bringing massive compute and memory onto a single piece of silicon and integrating it into a purpose- built system and software stack, we deliver exceptional AI speed for customers on premises and via the cloud. Our strategic partners and customers include hyperscalers, foundation model labs, AI-native and digital-native businesses, enterprises, and Sovereign AI initiatives. OpenAI, the world’s leading foundation model lab, selected us to be its fast inference solution. With Cerebras, OpenAI’s Codex-Spark users turn ideas into working software in seconds. This partnership is an example of tight hardware-software co-design with a leading frontier model lab. AWS, the world’s leading hyperscale cloud, has signed a binding term sheet with us to become the first hyperscaler to deploy Cerebras in its own data centers, providing massive distribution to a broad base of enterprise customers. Our customers use Cerebras solutions to run applications that demand speed, scale, and intelligence. This work includes training and serving large frontier models with near-instant responses, processing massive datasets in real time, and generating full-stack applications in a single step. Once customers adopt fast inference, user expectations for interactivity rise, and engineering teams shift from latency optimizations to other work, making it difficult to return to slower inference. We deliver our solutions to customers in several different ways. Organizations that require full data and infrastructure control can purchase Cerebras AI supercomputers for on-premises deployments. Customers seeking cloud flexibility can access Cerebras compute through consumption-based models on Cerebras Cloud or through partner clouds. For example, our high-speed inference services are available through partners, including AWS Marketplace, Microsoft Marketplace, IBM watsonx Model Gateway, Vercel AI Gateway, OpenRouter, and Hugging Face, enabling seamless adoption within existing workflows. Our ability to deliver differentiated performance has made us a strategic partner to many of our largest customers. Beyond providing compute infrastructure, we provide AI services to our customers to co-develop solutions to address their most complex challenges, from training state-of-the-art models to optimizing deployments for each application’s needs. These partnerships have expanded over time; notably, our top ten customers by year-to-
112
Table of Contents
date revenue through December 31, 2025 increased their aggregate spend with us by approximately 80% within 12 months of their initial purchase, often including contracts for co-development. AI is one of the fastest growing technologies in history. We believe that our high-speed AI solutions give us a meaningful competitive advantage in this market. We believe that further adoption of AI, accelerated by increased penetration, more frequent usage, and more complex applications, will continue to rapidly expand the market. According to IDC, investments in AI solutions and services are projected to yield a global cumulative impact of $22.3 trillion by 2030, representing approximately 3.7% of the global GDP. The combined market for AI training infrastructure and our addressable market within AI inference is estimated to be $251 billion in 2025 and is expected to grow to $672 billion by 2029—a 28% CAGR, according to Bloomberg Intelligence. This estimate indicates that AI inference will grow more than twice as fast as AI training infrastructure through 2029. With the fastest inference platform on the market, as benchmarked by Artificial Analysis, and a proven track record in large-scale training, we believe we are well-positioned to capture growth across both parts of the AI infrastructure market. Our growth reflects the broader acceleration of AI adoption. Our revenue increased from $24.6 million in 2022 to $78.7 million in 2023 and to $290.3 million in 2024, representing a more than tenfold increase over three years. Our revenue increased to $510.0 million in 2025, representing year-over-year growth of 76%. We earned net income of $237.8 million in 2025 and incurred net loss of $481.6 million in 2024. Our gross margin was 12%, 33%, 42%, and 39% in 2022, 2023, 2024, and 2025, respectively. We incurred non-GAAP net loss of $75.7 million in 2025 and $21.8 million in 2024, after excluding the impact of stock-based compensation expense and change in fair value (extinguishment) of forward contract liability from our GAAP net income (loss). For more information and for a reconciliation of non-GAAP net loss to net income (loss), see the section titled “Management’s Discussion and Analysis of Financial Condition and Results of Operations—Non-GAAP Financial Measures.” Industry Background AI is the Next Technological Shift Over the past 50 years, the compute industry has undergone a series of secular shifts, each of which expanded access to compute and transformed global productivity. In the 1990s, the Internet reshaped how people worked, communicated, transacted, and learned, catalyzing new industries and business models. In the 2000s and 2010s, the proliferation of mobile devices and the emergence of cloud computing delivered unprecedented flexibility, scale, and reach, supporting millions of new digital products and experiences. We believe AI represents the next major technological shift—one with the potential to exceed the transformational impact of prior cycles. In comparison to previous technology shifts, the adoption of AI is astonishing. Its market penetration has occurred multiple times faster than the PC and the cloud. ChatGPT reached 100 million users in less than 2.5
113
Table of Contents
months, more than twenty times faster than Facebook. As of September 2025, ChatGPT reported 700 million weekly active users. !busienss1ba.jpg
*busienss1ba.jpg*
According to Pew Research Center, as of June 2025, around 62% of U.S. adults interacted with AI at least several times a week, with 31% doing so almost constantly (at least several times a day), and one-third of U.S. adults under 30 saying they interacted with AI several times a day. Additionally, the Digital Education Council found in 2024 that 86% of higher-education students used AI. According to a McKinsey survey in 2025, the share of respondents saying their organizations are using AI in at least one business function has increased since their research last year: 88% reported regular AI use in at least one business function in 2025 compared with 78% a year ago. In the third quarter of 2025, Gallup reported daily use of AI in the workplace had more than doubled in the past 12 months, with 10% of U.S. employees reporting they used AI in their daily roles. The strong rate of AI adoption is driven by the simple fact that AI has transitioned from novelty to necessity and is now used across consumer and enterprise domains. Individuals and organizations rely on AI to solve problems, build products, accelerate research, improve patient outcomes, enhance decision-making, streamline operations, enable innovation, and deliver personalized experiences.
114
Table of Contents
The rise of AI depends on massive computational resources. This is where Cerebras fits in. !cerebras-drsx1219b.jpg
*cerebras-drsx1219b.jpg*
Inference is Driving the AI Compute Demand, as Frontier AI Models Grow More Capable AI is composed of two stages: training and inference. Training is the process of creating and teaching the AI model; inference is the process of using the model to generate responses. Early progress in AI came primarily from training larger models. Larger models, which used more compute during training, improved AI’s accuracy. In this training-centric era, inference was straightforward and required little computation; it simply generated answers from a trained model in a single step. Today, AI has entered a new era centered on inference. New techniques have emerged that make models smarter as they are being used. This approach—called “inference-time compute” or “test-time compute”—has become the dominant mode of inference. !cerebras-drsx1219c.jpg
*cerebras-drsx1219c.jpg*
115
Table of Contents
Instead of depending primarily on the trained model for accuracy, today’s frontier models—such as OpenAI’s GPT-5.4, Anthropic’s Claude Opus 4.7, and Google’s Gemini 3.1 Pro—perform substantial computation during inference to simulate reasoning. These models effectively “think through” the problem: planning steps, checking their own work, and refining responses before delivering a final, higher-quality result. These additional steps use substantially more compute during inference, while producing more accurate answers. !business4ba.jpg
*business4ba.jpg*
These reasoning capabilities have fundamentally changed how people use AI. Inference is no longer limited to answering questions; modern AI applications now perform actions on behalf of their users. They can directly book travel itineraries, code full web applications from scratch, help customers apply for mortgages, automatically analyze legal contracts for discrepancies, process insurance claims, and more. As a result, demand for AI inference has surged alongside the adoption of these smarter reasoning models that leverage more inference-time compute.
116
Table of Contents
Ultimately, inference compute demand is driven by the compounding effect of three forces: the number of users, the frequency of use, and the compute per use. Each of these forces is growing at an extraordinary rate, producing a geometric expansion of demand for inference and its underlying compute. !business5da.jpg
*business5da.jpg*
Reasoning during inference delivers smarter AI responses but requires significantly more compute. As models become more capable, users rely on them for increasingly ambitious tasks, further driving compute needs. Today’s workloads—including video generation, deep research, and long-form analysis—can require many orders of magnitude more compute than answering basic questions. Reasoning Makes Inference Speed a Necessity Speed enables reasoning models to deliver more accurate answers faster, reducing the frustration created by forcing customers to wait for answers. Reasoning changes the shape of inference. Reasoning systems do not complete tasks in a single request-and- response step. They execute a sequence of sequential and dependent steps—such as planning, refinement, and verification—until the task is completed. Each step consumes compute and contributes to total completion time. Slower execution of each step compounds, and then the task takes much longer to complete. Faster execution at each step shortens the overall time to answer. Complex tasks (harder problems) are more valuable to solve but they require the reasoning system to go through a longer sequence of steps. This amplifies the benefit of speed and the penalty for being slow. Speed enables more accurate answers to harder problems in less time. Speed expands the range of tasks that AI can address, thereby
117
Table of Contents
broadening its addressable market. Conversely, slow AI produces longer wait times, making many applications impractical to deploy. !business6da.jpg
*business6da.jpg*
Speed enables AI to address more complex, higher value tasks. This, in turn, brings new users to AI, who use AI more frequently and to solve more complex problems. And herein is the flywheel. More users, more frequent users, and more complex use cases all increase AI compute usage. Fast Inference Enables the Next Generation of AI Workloads, With Coding as a Clear Early Signal As AI uses more compute to tackle increasingly complex problems, a fundamental challenge emerges: everyone wants a better response for complicated requests, but nobody wants to wait to get a response. We are solving this problem. Cerebras Inference delivers answers up to 15 times faster than leading GPU-based solutions as benchmarked on leading open-source models. This speed advantage enables our solutions to deliver real-time performance for the most advanced reasoning models, enabling complex tasks to be completed more accurately and quickly.
118
Table of Contents
As discussed, fast and accurate results delight users, drive engagement, and unlock new classes of applications and business opportunities. Faster AI compute produces answers in less time, which drives more frequent usage, new types of applications, and therefore greater compute demand. These dynamics are already visible in the market. Three fast-growing categories—software development, deep research systems, and voice applications—illustrate the importance of speed. For these and many other similar applications, inference speed is a necessity. •AI-powered software development provides a clear early signal. Coding with AI is interactive and sensitive to delay. Delay impairs a developer’s train of thought, and as a result, developers are more likely to abandon tools that slow them down. AI can now write code. It reasons over large codebases and then uses the multi-step process previously described to generate, modify, and run code. Inference speed has become a primary determinant for adoption. Products such as Cursor, Claude Code, Codex, Windsurf, and GitHub Copilot act as autonomous collaborators—planning, editing, and validating code across repositories in response to natural-language instructions from developers. These systems require complex, multi-step tasks, including continuous reasoning and long-context memory. Fast inference is the only way to avoid frustrating wait times. AI-native coding products barely existed in 2023. Yet they collectively generated billions in ARR in 2025 and continue to accelerate. For example, AI coding applications like Lovable and Cursor are among some of the fastest growing developer tools in history. AI coding agents have become central to how software is written. Anthropic’s Claude Code is already at a reported annual revenue run rate of $2.5 billion as of February 2026; Claude Code’s creator said in January 2026 that he writes 100% of his code with AI. In addition, professional developers report that 42% of code is now AI-generated or assisted, according to a survey conducted by SonarSource in October 2025. By droves, software engineers are shifting from writing code to supervising fleets of AI coding agents. Faster inference means more productive engineers. Coding demonstrates a fundamental pattern in reasoning systems: wherever AI involves continuous interaction, multi-step reasoning, and sensitivity to response time, speed determines utility. Those same conditions are present across a growing set of AI applications. •Deep research systems apply similar reasoning to knowledge work, performing multi-step retrieval and synthesis across large datasets to deliver structured insights in real time. Platforms such as AlphaSense rely on real-time inference to sift through a higher volume of documents to help analysts and enterprises find answers faster. •Voice applications include conversational agents, avatars, and digital twins from companies like Meta, Tavus, and OpenCall. Real-time performance is critical for voice: sub-second latency makes interactions feel natural and gives these systems time to call tools or retrieve data mid-conversation for richer, contextual responses. Together, we believe these applications lead the way in the next phase of AI adoption: systems that think, act, and interact continuously, driving sustained demand for faster and more efficient compute infrastructure. In this environment, speed directly shapes usage. Long wait times limit real-time applications, stunt the diffusion of AI capabilities, and can inhibit new markets and applications. As a result, slow systems lose users, limit capability, and stall innovation, while faster systems are used more often and for more demanding workloads. We believe speed is a defining advantage in modern AI. Reasoning is intelligence, and intelligence compounds with speed. We believe the ability to deliver fast, scalable reasoning will define not only the next decade of technology, but also shape the future of how people work, create, and interact.
119
Table of Contents
Our Market Opportunity We address a large and rapidly growing market for AI infrastructure. According to Dell’Oro Group, worldwide data center infrastructure capital expenditures are expected to grow from $679 billion in 2025 to $1.7 trillion by 2030, representing a 21% CAGR. AI infrastructure increasingly dominates global IT spending. Training The AI training infrastructure market is expected to grow from approximately $185 billion in 2025 to $380 billion by 2029, a 20% CAGR, according to Bloomberg Intelligence. This market is characterized by large-scale capital buildouts as hyperscalers, foundation model labs, enterprises, and Sovereign AI initiatives, invest in developing foundation models and fine-tuning capabilities. We have demonstrated strong success in this market, most notably through hardware and models we’ve trained for G42, MBZUAI, GlaxoSmithKline, Sandia National Laboratory, the U.S. Department of Defense, and other training customers. Inference Based on Bloomberg Intelligence data, our addressable market within the AI inference market is expected to grow from approximately $66 billion in 2025 to $292 billion by 2029, a 45% CAGR. The AI inference market scales with the number of AI users, a number we expect to converge with the global internet user base over time. Inference compute can be accessed at the hardware level through on-premises deployments and at the cloud/API level, measured in tokens served. We serve both through Cerebras AI supercomputers, which are deployed directly in customer data centers, and Cerebras Inference Cloud, which addresses the token-based API market. The token-based market is expanding rapidly. In October 2025, Google reported Gemini was serving 1.3 quadrillion tokens per month—a market that was effectively zero before the launch of ChatGPT in late 2022. Cerebras Inference Cloud directly serves this market. Because the AI compute we provide is general purpose, we serve a wide range of models used across verticals—consumer applications, code generation, enterprise AI, and more. The same infrastructure that powers a chat application can power a financial model or a coding agent. The combined market for AI training infrastructure and our addressable market within AI inference is estimated to be $251 billion in 2025 and is expected to grow to $672 billion by 2029—a 28% CAGR, according to Bloomberg Intelligence. This estimate indicates that AI inference will grow more than twice as fast as AI training infrastructure through 2029, and we expect AI inference to represent an increasing share of total AI infrastructure demand as deployed models scale to serve global user bases. With the fastest inference platform on the market, as benchmarked by Artificial Analysis, and a proven track record in large-scale training, we believe we are well-positioned to capture growth across both markets. Our Solution We are building the fastest commercial AI infrastructure in the world. Our AI supercomputers are purpose built to make AI fast. They are built for the latency-sensitive, reasoning workloads that define modern AI. Our full-stack hardware and software platform is designed to complete AI tasks significantly faster and more efficiently than comparable GPU-based solutions, whether deployed on premises, through the Cerebras Cloud, or via partner clouds.
At the core of our solution is the Cerebras WSE, the largest and fastest AI processor ever brought to market in high volumes. The WSE combines 900,000 compute cores, 44 gigabytes of on-chip memory, and 21 petabytes of memory bandwidth on the largest commercial chip ever built. The WSE-3 is 58 times larger than NVIDIA’s B200 chip. The WSE has 19 times more transistors, 250 times more on-chip memory, and 2,625 times more memory bandwidth than NVIDIA’s B200 package, which contains two individual chips.
120
Table of Contents
Each WSE is housed inside a Cerebras CS-3 system, our fully integrated AI compute system that includes advanced cooling, power delivery, and interconnect technology. Multiple CS-3 systems connect to form Cerebras AI supercomputers deployed on premises in customer data centers and in the cloud.
*cerebras-drsx12191a.jpg*
122
Table Of Contents
Our software platform makes wafer-scale computing simple to use. It spans the full AI life cycle—from model programming and compilation, to training and inference, to cluster orchestration. •Cerebras Compiler compiles PyTorch models directly to the WSE, eliminating the need for CUDA or distributed programming and providing an easy-to-use developer experience. •Cerebras Inference Serving Stack delivers ultra-low-latency inference with industry-standard APIs for production use. •Cerebras Cluster Manager orchestrates multiple CS-3 systems into one logical AI supercomputer, handling scheduling, telemetry, and health monitoring at scale. Because every layer is co-designed with our hardware, customers can scale training and inference across frontier-size models without rewriting code or managing distributed infrastructure.
Our technology is designed to be delivered in the form that best accelerates a customer’s AI roadmap. Our platform is designed for flexibility—meeting organizations where they are, and scaling with them as their ambitions grow. •Cerebras Cloud: Provides high-performance AI compute through a simple API, allowing customers to serve open-source, fine-tuned, or proprietary models with production-grade reliability. •Partner Clouds: Offer seamless access to Cerebras systems through leading cloud providers including AWS Marketplace, Microsoft Marketplace, IBM watsonx Model Gateway, Vercel AI Gateway, OpenRouter, and Hugging Face, extending our reach across the global AI ecosystem. •On-Premises Deployments: Deliver fully integrated AI supercomputers and install them directly in customer environments, giving enterprises, Sovereign AI initiatives, national laboratories, and defense organizations complete control over data, performance, and operations. We also operate and manage large clusters of AI supercomputers for some of our customers. •Hybrid Deployments: Enable customers to move fluidly between on-premises and cloud environments through a unified software stack, maintaining consistent performance and workflows as they scale.
123
Table Of Contents
Customers choose the consumption model that fits their needs—buying inference by the token, running training workloads by the week or month, reserving dedicated capacity for long-term production deployments, or purchasing on-premises infrastructure. !business15ca.jpg
*business15ca.jpg*
Our AI experts accelerate customers’ ability to take AI applications from concept to production. With deep experience training and deploying frontier-scale models across modalities, our team helps customers select model architectures, prepare large-scale training data, and train and fine-tune models for production. We also design optimized deployments for customers—training draft or speculative decoding models and tuning configurations to balance latency, throughput, and cost for each application. We excel at turning AI ambition into business results. By augmenting customer teams with advanced AI expertise, we help customers design, build, and deploy custom models that often outperform existing state of the art, giving customers a meaningful competitive advantage. Together, our hardware platform, unified software, and AI model services form an integrated platform that becomes increasingly valuable over time. As customers build models, workflows, and applications on Cerebras, the platform can become deeply embedded in their AI development and operations, leading to durable relationships. What This Means for Customers Our customers, which include hyperscalers, foundation model labs, AI-native and digital-native businesses, enterprises, and leaders of Sovereign AI initiatives, complete tasks dramatically faster than on GPU-based systems. Faster reasoning improves user experience, increases engagement, accelerates iteration, and enables new classes of AI applications. This speed advantage compounds in production environments, where reduced latency and shorter training cycles have meaningful business impact. Key Customer Benefits Through our full-stack AI offerings, we deliver tangible improvements across four key dimensions that define AI value in the real world: speed, quality, cost, and simplicity.
124
Table Of Contents
Our systems achieve dramatically faster inference than GPU clusters, enabling applications such as real-time coding agents, nearly instant deep research, and digital twins that were previously impractical or impossible. Customers describe the leap in inference speed as akin to going from dial-up to broadband—an advancement that redefines what AI can do. New classes of products that customers have built and use daily with Cerebras include: •Real-time coding agents: Copilots that read, write, and debug code nearly instantly—turning AI into an interactive programming partner. •Nearly instant deep research agents: Systems that analyze thousands of documents in seconds, accelerating market, scientific, and policy research. •Digital twins: Lifelike AI personas that think, speak, and react in real time. With Cerebras, avatars respond without awkward delays, and carry conversations that are more natural and interactive.
On GPUs, latency forces a tradeoff between speed and intelligence. Developers often have to limit the accuracy of a response in order to have it delivered in a reasonable amount of time. Our offerings are designed to remove this tradeoff. Customers run long-context, multi-step reasoning models interactively, delivering higher-quality results without comparable delays. Our inference speed allows developers to use substantially more reasoning tokens while maintaining the same end-to-end task completion time. We turn quality from a limitation into a feature; customers can now serve some of the largest models at full strength, in nearly real time.
Moving data from one chip to another is one of the most power-intensive parts of AI compute. And power is the largest contributor to operating expenses in AI compute. Our wafer-scale architecture keeps data on-chip, reducing data movement significantly, which in turn reduces power consumption. It also eliminates layers of costly and complex networking equipment. By way of comparison, moving a bit of data on the WSE-3 consumes a fraction of the energy required to move the same bit of data over GPU interconnects. Because our performance advantages stem from fundamental architectural efficiency, we expect these benefits to endure across future generations that continue to build on our wafer-scale technology.
We eliminate the complexity of distributed programming across GPU clusters, which is one of the most challenging aspects of AI deployment. Even extremely large models run without code changes, and scale automatically and seamlessly across clusters of Cerebras systems. Because training, fine-tuning, and inference all occur on a unified platform, customers avoid the operational overhead of moving between different compute environments, enabling inference, fine-tuning, and training from scratch on the same cluster. Cerebras Compiler’s PyTorch integration makes model customization and compilation simple, the Inference Serving Stack enables deployment of frontier-sized models in minutes, and our AI experts support customers throughout the model life cycle to accelerate results. Cerebras’s deployment platform also allows customers to run models that were not trained on Cerebras hardware and still achieve exceptional inference performance.
125
Table Of Contents
Factors Preventing GPUs From Being Faster at AI AI inference speed is limited by how fast data moves between memory and compute; this is called memory bandwidth. When a large language model generates a response, it predicts one word (token) at a time. Each token generated requires a large amount of data—all of the model weights—to be moved from memory to compute. Because each token depends on the previous one, this work cannot be parallelized, making AI inference speed fundamentally limited by memory bandwidth. !business16aa.jpg
*business16aa.jpg*
GPUs were designed for graphics workloads. Graphics can tolerate slower memory movement because the highly parallelizable workload allows the GPU to keep many compute cores busy while more data is moved over, masking the memory latency. These workload characteristics made high-capacity off-chip memory, placed far away from the compute processor, a strong architectural choice for graphics. But the tradeoff of off-chip memory is speed. Off-chip memory connects to the compute processors through a narrow data “pipe” with low memory bandwidth. This was the right tradeoff for graphics, but it creates a critical limitation for AI speed, where data movement is the bottleneck. The result is a GPU “memory wall” for AI, where a GPU’s memory bandwidth cannot keep up with compute for AI workloads, and thereby limits the speed with which AI answers can be generated. These are not software issues. They are the physical limits of the memory + GPU architecture. In order to be fast, we believe AI requires a fundamentally different architecture that solves the memory wall. Such architecture must provide vastly more memory bandwidth to enable data to move more quickly between memory and compute, which can thereby accelerate the generation of AI responses.
126
Table Of Contents
Our Technology Wafer-Scale Integration: The Foundation Cerebras started with a simple question: How could a new class of processors be designed with the singular goal of solving the compute challenges presented by AI? Beginning with a clean slate, how could we avoid the trade-offs made for graphics and other workloads to ensure that every transistor, every single part of the processor, was optimized for the requirements of AI? Our answer is wafer-scale integration. Wafer-scale integration enabled us to use a vastly faster memory and avoid the complexity of switches and routers and associated complexity necessary to link together thousands of GPUs. SRAM is the fastest memory to date. But existing industry players could not use as much SRAM because they could not fit it on their chip. By building a chip 58 times larger than NVIDIA’s B200 chip, we can maximize fast, on-chip SRAM and get the benefits of two worlds: (1) significantly more memory capacity because we built such a big chip, and (2) the benefits of the massive bandwidth provided by SRAM. Wafer scale enables us to deliver a solution with 2,625 times more memory bandwidth than NVIDIA’s B200 package, which is how we are able to deliver inference at extremely fast speeds. The second fundamental advantage provided by wafer-scale integration is that it kept the wafer intact. Instead of building a wafer, cutting it into dozens of small GPUs, and using expensive, power-hungry switches, and complex cables to wire them back together, our solution consists of one processor that is the size of an entire silicon wafer. This reduced the need, cost, managerial complexity, and power draw of much of the networking stack required to build a GPU solution. Our wafer-scale solution unifies compute and memory and communications on the same piece of silicon, eliminating the data-movement bottlenecks that slow GPU systems. The Underpinnings of Wafer-Scale Integration We solved a problem that flummoxed the compute industry for its entire history: how to build chips the size of full silicon wafers. The advantages of size were well known. But no company had ever brought a wafer-scale solution to market. To make wafer-scale commercially viable, we invented and productized two foundational semiconductor technologies: •Multi-die interconnect: Traditionally, die—regions of silicon containing an integrated circuit—are individually stamped onto a silicon wafer and then cut up (“diced”) into small, separate chips. Prior to Cerebras, the largest known chip was about 840 mm. We invented technology to interconnect these otherwise independent die together at the wafer level, at the semiconductor fabrication plant. The inter-die connectivity uses a proprietary cross-reticle connection that is integrated into our overall fabrication process. This allowed us to use existing processes to do something we believe had never been done before —namely, deliver a wafer that communicated across the entire 46,225 mm of silicon and therefore is a single massive processor. •Fault-tolerant architecture: A primary factor in the commercial viability of a semiconductor is the yield. Flaws are present in wafers. Large chips have a higher probability of hitting such a flaw. Traditionally, chips with flaws have been thrown out or “down binned,” that is, sold as a less capable part. Thus, using traditional techniques, larger chips have lower yield and are therefore more expensive. We designed the architecture to absorb and route around defects using redundant building blocks—similar to a hyperscale data center but on the wafer. Flaws are designed to be recognized, shut down, and routed around. Redundant building blocks are used to re-form a logically functional whole. This approach had been
127
Table Of Contents
previously used in memory manufacturing to achieve near-perfect yield, but to our knowledge, prior to Cerebras had not been used to build processors. These innovations made wafer-scale computing commercially viable for the first time in semiconductor history. !business17ga.jpg
*business17ga.jpg*
The Cerebras Chip, System, and Software Cerebras delivers a full-stack AI infrastructure solution. It contains innovations at each layer. At the base is the Cerebras WSE, our wafer-scale processor. Each WSE is integrated into a CS-3 system with advanced power delivery, cooling, and system management. Multiple CS-3 systems link together to form Cerebras AI supercomputers that are deployed in data centers around the world. Lightweight management and orchestration software operate these systems as one logical computer, while our training and inference platforms make it simple to run large models at scale. Because each layer is designed with the
128
Table Of Contents
others in mind, the platform delivers consistent performance, reduced infrastructure complexity, and faster time to deployment and results.
At the heart of our platform is the Cerebras WSE, the world’s largest and fastest commercialized AI processor. A single WSE replaces an entire cluster of GPUs by combining 900,000 compute cores and 44 gigabytes of on-chip memory on one piece of silicon, with 21 petabytes per second of on-chip memory bandwidth. The WSE-3 is 58 times larger than NVIDIA’s B200 chip. The WSE-3 also has 19 times more transistors, 250 times more on-chip memory, and 2,625 times more memory bandwidth than NVIDIA’s B200 package, which contains two individual chips. We believe our architecture solves for memory bandwidth, which is a primary bottleneck in modern AI. By keeping compute and memory on a single chip, WSE-3 eliminates the off-chip data transfers that dominate GPU latency and power consumption. As a result, our systems are faster, simpler to program, and more power-efficient than GPUs on AI tasks. Fast inference depends on memory bandwidth. Below, we show the traditional GPU architecture with HBM, a type of off chip DRAM, and a GPU. For the GPU to generate a single word based on an inference prompt for a 70 billion parameter model, it must move more than 140 gigabytes of data from memory to compute. That is roughly 100 1-hour HD movies. This is to generate a single word. And this must be done again and again for each word in sequence. !business4aa.jpg
*business4aa.jpg*
The speed of generating a response is limited by the rate at which data can move from memory to computer. In the figure below, we show the underpinning of our performance advantage. We have a 2,625 times larger pipe
129
Table Of Contents
between memory and compute. More data can move though our pipes, meaning we generate “words” (tokens) much more quickly. !business5ba.jpg
*business5ba.jpg*
The WSE-3 is deployed inside the CS-3 system, a data center-ready appliance engineered to support wafer- scale operation and integrate seamlessly into enterprise and Sovereign AI environments. The CS-3 provides the power delivery, cooling, networking, and system management required to operate a wafer-scale processor reliably and at scale. Multiple CS-3 systems can be connected to form Cerebras AI supercomputers, which function as a single logical computer for large-scale training and inference.
131
Table Of Contents
Our software platform extends our hardware advantage by making wafer-scale computing simple to use and highly efficient. Our software spans the full AI life cycle—from programming and compiling models, to training and inference, to orchestration across large clusters. Each layer is co-designed with our hardware to deliver maximum performance with minimal developer effort. Model Programming and Compilation. Our Cerebras Compiler (CSoft) makes it simple to run large language models on our systems. CSoft is core to our solution and provides intuitive usability for developers. CSoft eliminates the need for low-level programming in CUDA or other hardware-specific languages. For both training and inference, our CSoft platform enables developers to easily represent and map large language models onto the Cerebras Wafer-Scale Engine using familiar frameworks such as PyTorch. Starting from a user’s PyTorch model, the CSoft graph compiler automatically maps model operations to the WSE, creating an optimized executable without user-level intervention. CSoft allows machine-learning users to accelerate training and inference on models of any size, scaled across any configuration of the Cerebras AI supercomputer, just by changing one number in a configuration file, simulating a single-device programming experience without the complexities of distributed programming. This drastically reduces operational overhead and speeds up developer iteration time and business impact. Inference Serving Stack. Our Cerebras Inference Serving Stack manages model hosting, scaling, and request routing across Cerebras systems and clusters. It provides real-time observability and load balancing, enabling ultra-low-latency inference for production workloads. Customers can serve both open-source and proprietary models through standard APIs, including industry-standard endpoints, with consistent performance across on-premises and cloud deployments. Orchestration and Life Cycle Management. Our Cerebras Cluster Manager orchestration software unifies multiple CS-3 systems into a single logical computer, managing scheduling, telemetry, and health monitoring. Built- in observability of all hardware and software components is designed to ensure reliability and high utilization across on-premises and in cloud environments. This orchestration layer also allows customers to switch seamlessly between training and inference on the same systems. With simple commands, CS-3 systems can be reconfigured from large-scale model training to real-time inference, driving utilization and shortening deployment cycles. Together, these components form a unified software platform that integrates seamlessly with our hardware to deliver a complete, end-to-end AI computing system that can be deployed on customer premises or in the cloud.
132
Table Of Contents
Because our software and hardware are co-designed, customers can train and/or deploy frontier-scale models with consistent and simple workflows—without rewriting code or managing distributed infrastructure. !business1ba.jpg
*business1ba.jpg*
Technology and Roadmap Wafer-scale integration is not a single achievement—it is a collection of technologies and processes with a multi-generation roadmap. Each successive WSE generation (from 16 nanometer to 7 nanometer and now to 5 nanometer) has delivered substantial improvements in performance, memory bandwidth, efficiency, yield, and manufacturability, without requiring changes to how developers program or deploy models. Competing approaches—such as multi-die packages and chiplet based designs—remain constrained by the physics of small chips and limited off-chip memory bandwidth. Even with advances in packaging technology, these architectures cannot match the bandwidth, locality, or simplicity of computation that result from keeping compute and memory together on a single piece of silicon. Our roadmap builds on the advantages of wafer-scale integration. We intend to invest heavily in research and development to continue to expand on-chip memory and memory bandwidth, improve interconnect density, and leverage advancements in process technology to increase transistor counts and reduce power in future WSE generations. As a result, we expect that future generations of WSEs will have faster compute, and more and faster memory and communication onto and off of the wafer. Because the WSE presents itself as a single programmable device, these improvements compound naturally in both performance and simplicity, without introducing the complexity of massive distributed compute clusters of GPU solutions. The same architectural foundation also supports long-term extensibility across emerging AI workloads. As models grow in size, increase in reasoning depth, and shift toward real-time, multi-step interactions, they place even greater emphasis on memory bandwidth and locality—all areas where wafer-scale architectures possess inherent, structural advantages.
133
Table Of Contents
Our roadmap includes development of a disaggregated inference-serving solution. Inference disaggregation is a technique that separates AI inference into two stages: prompt processing, or “prefill,” and output generation, or “decode.” These two stages have different computational characteristics. Prefill is natively parallel and requires very little memory bandwidth. Decode, on the other hand, is inherently serial and memory bandwidth intensive. Decode is typically the bottleneck. It dominates total inference time, and defines the speed of the user experience. Cerebras’s wafer-scale engine would be the fastest at both prefill and decode, but in relative terms, it is much faster at decode. Our wafer-scale architecture and ultra-high memory bandwidth delivers faster output token generation where speed matters most. Disaggregated inference would allow Cerebras to operate alongside other architectures, serving as the high-performance engine for decode while other systems handle prefill. We believe wafer-scale computing positions us as a leader in AI infrastructure, providing a long-term technology roadmap designed to scale with the requirements of modern and future AI systems. Competitive Strengths 1.Our culture of fearless engineering has enabled us to do pioneering engineering work; we are the only company ever to deliver a wafer-scale processor to market. Our culture of fearless engineering enables us to solve problems that others failed to solve or were afraid to tackle. As a result, we have solved problems that had remained unsolved for the entire 75-year history of the compute industry, namely wafer- scale integration. A culture of fearless engineering is a foundation for our continued innovation. 2.We have durable advantages rooted in our unique silicon architecture. We believe wafer-scale integration is a fundamental advantage in AI compute, enabling large amounts of high-speed memory and hundreds of thousands of compute cores to reside close together on the same piece of silicon. We have now delivered three generations of wafer-scale processors at the 16, 7, and 5 nanometer nodes. We believe these will be the foundation of our future generations of silicon. 3.We are an end-to-end systems company. From inception, we co-designed our wafer-scale engine, our CS-system, and our software stack for optimal AI performance. We were among the first in the AI community to deliver water cooling to the processor, enabling us to run colder and extend our processors’ lifetime. The co-design of processor, system, and software is a meaningful competitive advantage. 4.We are building the fastest inference infrastructure in the world. On Cerebras infrastructure, AI responses are up to 15 times faster than leading GPU-based solutions as benchmarked on leading open- source models. Third-party benchmarker Artificial Analysis wrote in August 2024, “Cerebras Inference is achieving the fastest speeds we have ever benchmarked on Artificial Analysis.” Speed is customer experience. It enables more accurate answers in less time. It enables applications that require real-time interaction such as coding agents, research agents, and voice interfaces. Speed changes the way companies design their experiences; it changes team structures and behaviors; it changes expectations and the perception of what is possible, which can make returning to slower speeds more painful. 5.We are serving some of the largest and most demanding customers in the AI market. We are engaged with customers such as OpenAI, the world’s leading foundation model lab, and AWS, the world’s leading hyperscale cloud, who have stringent requirements for performance, scale, and reliability. We offer a full- stack hardware and software platform that can be optimized for each customer’s workloads and paired with AI services in order to deploy and operate high-capacity, production-grade systems without requiring customers to manage complex infrastructure. 6.We operate at massive scale with more than 100 exaflops of deployed compute. In collaboration with our partners, we have trained some of the largest models in the industry, gaining unique experience and providing rare insight. We license space in six data centers in North America, providing geographic
134
Table Of Contents
redundancy and regional deployment options for customers with data residency or network time requirements. !business2f.jpg
*business2f.jpg*
*business3d.jpg*
135
Table Of Contents
*business4d.jpg*
Our Business Model: Make Buying Easy, by Reaching Customers Where They Are The AI market is one of the fastest-growing technology sectors in history. Within this rapidly evolving landscape, we engage customers through a combination of direct sales and an ecosystem of strategic partners. Our sales organization, together with our partners’ sales teams, delivers our high-performance AI solutions through multiple consumption models: (i) on premises, (ii) through our own Cerebras Cloud, (iii) via partners’ clouds, or (iv) through hybrid combinations of these approaches. This flexible delivery model allows customers to adopt our technology in the manner that best aligns with their procurement preferences, operational requirements, and infrastructure strategies. Our product portfolio spans on-premises AI supercomputing systems, cloud-based compute for training and inference, and forward-deployed AI services to help customers accelerate the creation and deployment of AI capabilities. On-Premises Solutions Cerebras AI supercomputers support both model training and inference and are deployed directly within a customer’s environment. This deployment model is well suited for customers with regulated and high-security environments that require full control over data, infrastructure, and system behavior. Our on-premises customers include large enterprises, national laboratories, the U.S. Department of Defense, and Sovereign AI initiatives. Commercial Model for On-Premises Deployments On-premises customers procure our AI supercomputers through a traditional purchase-order process with payment received upon delivery or acceptance. Each system combines tightly integrated hardware and software and the purchase includes a separate renewable software subscription for continuous updates and upgrades, generating a
136
Table Of Contents
recurring revenue stream. On-premises deployments are often paired with our forward-deployed AI services, in which we assist customers with data preparation, model architecture design, training management, inference optimization, and, in select cases, ongoing system operations. Cloud Solutions We also provide access to our high-performance compute through Cerebras Cloud and through our partner cloud platforms, which include AWS Marketplace, Microsoft Marketplace, IBM watsonx Model Gateway, Vercel AI Gateway, OpenRouter, and Hugging Face. These offerings enable customers to utilize the full capabilities of our AI supercomputers without incurring the capital expenditures associated with building or maintaining on-premises infrastructure, and without the operational complexity of assembling and managing training or inference software stacks. Provisioning is highly streamlined, allowing customers to begin using our cloud resources within minutes. Cerebras Cloud serves a broad spectrum of users—from individual developers to some of the world’s largest enterprises. Customers run open-source, fine-tuned, and proprietary models for both training and inference workloads. Across all use cases, our cloud offerings provide access to ultra-high-performance AI compute. Customers procure cloud capacity from us and our cloud partners through two primary models: Dedicated Capacity and On-Demand. Dedicated Cloud Capacity Customers can contract for dedicated AI compute capacity for training or inference over defined terms. These contracts are generally structured as take-or-pay commitments, under which customers pay for dedicated compute capacity irrespective of utilization. Dedicated capacity provides availability and is well suited for production deployments and large-scale workloads. Customers in this model include leading hyperscalers, foundation model labs, AI-native and digital- native businesses, enterprises, and Sovereign AI initiatives operating open-source, fine-tuned, or proprietary models. Dedicated capacity contracts also include access to tailored workload telemetry that enables customers to optimize performance on our systems. This deep integration supports long-term engagement and increases platform stickiness. On-Demand Cloud Capacity For customers with variable or unpredictable workload requirements, we offer a consumption-based “pay-as- you-go” option. In this model, customers either purchase tokens—which represent units of compute—as they consume them, or pre-purchase token bundles and draw down their balance as workloads run. Enterprise customers are billed monthly, while individual developers access the service through a self-service portal. The on-demand model allows customers to scale elastically and is particularly effective for dynamic inference workloads. Historically, many customers have begun with on-demand usage and transitioned to dedicated capacity as their workloads expand. Customers and Go-to-Market Strategy Our customers include many of the world’s leading AI organizations. These span frontier model developers; hyperscalers; AI-native companies; as well as enterprises, research institutions, and national laboratories. Across these segments, customers rely on our solutions to accelerate model development and to deploy AI capabilities at production scale. We go to market through a combination of strategic partnerships, direct sales, channel partnerships, and product-led expansion.
137
Table Of Contents
Strategic Partnerships We partner with frontier model labs and hyperscalers to co-develop and deploy AI systems at scale alongside some of the most influential players in the AI ecosystem. OpenAI. We signed the MRA with OpenAI on December 24, 2025. On January 23, 2026, we began delivering capacity to OpenAI, and on February 12, 2026, OpenAI’s Codex-Spark model, powered by Cerebras infrastructure, was made available to the public. Spark is OpenAI’s model designed for real-time coding. Using Cerebras, OpenAI’s customers can translate ideas into working software in seconds, enabling developers to create software at the speed of thought. OpenAI has committed to purchase 750MW of Cerebras inference compute capacity over the next three years. Our partnership with OpenAI also allows for collaboration and co-design across both frontier model development and hardware architecture. This hardware-software co-development enables OpenAI to design models built for our hardware architecture and Cerebras to evolve hardware design in response to the needs of upcoming frontier model architectures. This creates a continuous feedback loop that can help our systems prepare for the next generation of AI, establishing a structural advantage for Cerebras. AWS. We signed a binding term sheet with Amazon Web Services for AWS to become the first hyperscaler to deploy Cerebras systems in its data centers. Deployment in AWS data centers will require us to meet strict standards for performance, scale, and reliability. Pursuant to the term sheet, we will create a co-designed, disaggregated inference-serving solution that will integrate AWS Trainium3 chips with Cerebras CS-3 systems, connected via high-bandwidth networking, to partition inference workloads across Trainium3 and CS-3. Each system will perform the type of computation at which it most excels. The approach is expected to deliver 5 times more token throughput in the same hardware footprint, at up to 15 times faster speeds compared to leading GPU-based solutions as benchmarked on leading open-source models. Direct Sales, Channel Partnerships, and Product-Led Expansion Direct Sales. We employ a targeted named-account strategy built on deep technical and commercial engagement. Dedicated account teams work closely with customer executives and their engineering leadership to identify business-critical workloads and then successfully integrate, optimize performance, and scale the deployment of Cerebras systems. This hands-on engagement builds operational trust and frequently results in the expansion of initial projects into multi-system or multi-year commitments. Alongside our direct sales force, we maintain a dedicated team of AI experts. This team provides customers with access to leading AI expertise, so that they are positioned to leverage our technology effectively. Channel and Technology Partnerships. To broaden our market reach, we leverage a diversified network of channel and technology partners. Cerebras solutions are accessible through AWS Marketplace, Microsoft Marketplace, IBM watsonx Model Gateway, Vercel AI Gateway, OpenRouter, and Hugging Face, allowing developers and enterprises to incorporate Cerebras performance seamlessly into existing workflows and deployment environments. Product-Led Growth. Our product-led growth motion introduces developers, startups, and emerging AI organizations to Cerebras through our self-serve inference platform and API. These early interactions often seed future named-account relationships. In addition, partnerships with cloud providers, system integrators, and software platforms that embed Cerebras capabilities into established workflows further expand access and reinforce our enterprise sales motion. We further extend our presence through integrations with widely used open-source development environments —including Visual Studio Code, Cline, RooCode, and OpenCode—embedding our technology, powered by the
138
Table Of Contents
Cerebras Cloud, directly where developers build and iterate. These channels enhance visibility within software- developer communities, foster product-led adoption of our self-serve offerings, and extend the reach of our platform beyond traditional enterprise sales.
*customerspotlight1ca.jpg*
*customerspotlight2e.jpg*
*customerspotlight4f.jpg*
*customerspotlight5g.jpg*
143
Table o f Contents
Sales and Marketing Our sales and marketing strategy centers on deep market understanding and customer-centric product development. We leverage our extensive market knowledge, proven track record in delivering large-scale compute solutions, and close customer collaborations to optimize our product roadmap. This is designed to ensure our solutions consistently deliver significant value to our customers. We focus our sales and marketing efforts on industry leaders, specifically large enterprises domestically and abroad with rich data assets. Our customers are seeking to leverage their rich proprietary data and combine it with Cerebras’s industry leading compute and AI expertise to build a durable competitive advantage. Among our customers, word of success travels quickly, and as a result, it is very important to our future that we maintain strong and collaborative relationships and that we invest in the success of our customers. We utilize master purchase agreements, purchase orders, and statements of work, to define work scope, price, quantities, delivery terms, warranties, and software subscriptions. We predominantly sell our solutions directly to customers via on-premises hardware or via the cloud, based on a dedicated capacity or consumption-based model. Research and Development We are committed to relentless innovation in both hardware and software to address the rapidly-evolving computational needs of AI. We dedicate significant resources to ongoing research and development. We invest heavily in attracting and retaining a global team of highly skilled engineers across dedicated facilities in the United States, Canada, and India. This unwavering commitment to innovation fuels our growth and positions us as a leader in the AI landscape. Manufacturing and Suppliers We operate a fabless manufacturing model, strategically partnering with industry leaders for the production of our AI compute systems which include ICs, boards, and systems. Our core manufacturing partners include: TSMC, a leading semiconductor foundry, fabricates our cutting-edge WSEs. Advanced Semiconductor Engineering (“ASE”) handles specialized processes, including the deposition of redistribution layers, and we manage final wafer packaging, assembly, and testing in our Sunnyvale, California facility. We also use a small number of third parties to manufacture subassemblies and critical components such as printed circuit boards, I/O subsystems, cooling assemblies and power delivery modules. The manufacturing process is subject to extensive testing and verification. Our supply chain is designed for flexibility and for quality, as we plan to ramp up production to meet the growing global demand for our AI compute systems. Simultaneously, we are committed to rigorous quality control throughout the manufacturing process to confirm reliability in even the most demanding environments at our customer facilities. Our contract manufacturing partners perform system assembly and extensive testing, and we have verification protocols in place at every stage including post assembly. Final system-level burn-in and test is conducted by Cerebras. Our quality processes include high production test coverage, full product traceability, and extensive post assembly burn-in. We employ a dedicated quality team that continuously monitors feedback during manufacturing and after deployment. This data-driven approach allows us to improve our product quality and reliability, and enables us to meet the stringent demands of our customers worldwide. Intellectual Property Protecting our intellectual property and proprietary technology, including our AI products and solutions, is an important aspect of our business. We rely on a combination of intellectual property rights, including patent, trademark, trade secret, and other related laws in the United States and internationally as well as confidentiality procedures and contractual provisions to protect, maintain, and enforce our proprietary technology, intellectual
144
Table o f Contents
property rights, and brand. Our intellectual property portfolio includes patents, trademarks, proprietary software, and trade secrets. As of March 31, 2026, we owned 96 issued patents and 50 pending patent applications globally. Of these, 50 are issued U.S. patents and 47 are pending U.S. patent applications. Our issued patents and pending patent applications generally relate to the design and fabrication of large-scale (e.g., wafer scale) processors, the assembly, packaging, and cooling of processors, and hardware, and software architectures for accelerated deep learning and for inference. The expiration dates of the U.S.-issued patents are between 2038 and 2041, not taking into account any applicable patent term extensions. We routinely review our development efforts to assess the existence and patentability of new inventions. We have a policy of requiring employees and consultants to execute confidentiality agreements upon the commencement of an employment or consulting relationship with us. Our employee and independent contractor agreements also require relevant employees and independent contractors to assign to us all rights to any inventions made or conceived during their employment or engagement with us. In addition, we typically require individuals and entities with whom we discuss potential business relationships to sign non-disclosure agreements that contain customary confidentiality provisions. Competition We offer a purpose-built AI compute platform. Our hardware primarily competes against solutions from NVIDIA Corporation, Advanced Micro Devices, Inc., Intel Corporation, as well as AI accelerators developed by hyperscalers and private companies. We also compete against full-service cloud service providers such as Amazon.com, Inc. (AWS), Microsoft Corporation (Azure), Alphabet Inc. (Google Cloud Platform), and Oracle Corporation, as well as AI-optimized specialized clouds such as CoreWeave, Inc. and other neo-clouds. We believe that our ability to remain competitive will depend on how well we are able to anticipate the features and functions that customers will require and whether we are able to deliver consistent volumes of our products and services at acceptable levels of quality and at competitive prices. We expect competition to increase from both existing competitors and new market entrants with products that may be lower priced than ours or may provide better performance or additional features not provided by our products and services. In addition, it is possible that new competitors or alliances among competitors could emerge and acquire significant market share. Some of our competitors have greater marketing, financial, distribution, and manufacturing resources than we do and may be more able to adapt to customers or technological changes. We expect an increasingly competitive environment in the future. Human Capital As of December 31, 2025, we had 708 employees, including 426 in the United States, and we have employees located internationally, including in Canada and India. We maintain a full-time workforce and supplement our workforce with contractors and consultants. To our knowledge, none of our employees are represented by a labor union or party to a collective bargaining agreement. We consider our relationships with our employees to be good. Our human capital resources objectives include, as applicable, identifying, recruiting, retaining, incentivizing, and integrating our existing and new employees. The principal purposes of our equity incentive plans are to attract, retain, and reward personnel through the granting of stock-based compensation awards in order to motivate these individuals to perform to the best of their abilities, enabling us to achieve our objectives. Facilities Our corporate headquarters is located in Sunnyvale, California, where we lease approximately 68,000 square feet, pursuant to a lease agreement that expires in November 2027, subject to the terms thereof. We lease additional facilities in Canada and India for research and development.
145
Table o f Contents
We also enter into agreements for offsite colocation facilities to house and operate our AI supercomputers. We enter into these agreements for our own corporate purposes as well as on behalf of our customers. Currently, these data center facilities are in California, Oklahoma, and Canada. We believe that our facilities are suitable to meet our current needs. We intend to expand our facilities or add new facilities as we grow, and we believe that suitable additional or alternative spaces will be available on commercially reasonable terms, if required. Government Regulations We are subject to many U.S. federal and state laws, rules, and regulations, as well as laws, rules, and regulations imposed by various non-U.S. governmental authorities, including those related to AI, intellectual property, tax, import and export requirements, anti-corruption, economic and trade sanctions, national security and foreign investment, foreign exchange controls and cash repatriation restrictions, data privacy and security requirements, competition, advertising, employment, product regulations, environment, health, and safety requirements, and consumer laws. These laws and regulations are complex, are constantly evolving, and may be interpreted, applied, created, or amended, in a manner that could harm our business. The import and export of our offerings are subject to laws and regulations, including international treaties, U.S. and various non-U.S. export controls and sanctions laws, customs regulations, and other trade rules. The scope, nature, and severity of such controls varies widely across different countries and may change frequently over time. Such laws, rules, and regulations may delay the introduction of some of our offerings or impact our competitiveness through restricting our ability to do business in certain countries or territories or with certain parties (including certain governments) or certain jurisdictions. U.S. export restrictions also require us to obtain licenses from the U.S. Department of Commerce to allow the export or transfer of our offerings, and there can be no assurance that export permissions will be granted. These restrictive governmental actions and any similar measures that may be imposed on U.S. companies by other governments could limit our ability to conduct business globally. See the section titled “Risk Factors” for additional information regarding risks we face related to government regulation. Legal Proceedings From time to time, we may be subject to legal proceedings, claims, and investigations in the ordinary course of business. We are not presently a party to any litigation to which the outcome, we believe, if determined adversely to us, would individually or taken together have a material adverse effect on us. We cannot predict the results of any such proceedings, claims, or investigations, and despite the potential outcomes, the existence thereof may have a material adverse impact on us due to diversion of management time and attention as well as the financial costs related to resolving such matters.
146
Table o f Contents