← Overview

Business

16,000 tokens · 85,001 chars

Business


BUSINESS Our Mission We believe AI is the most transformative technology of our generation. Our mission is to accelerate AI by making it faster, easier to use, and more energy efficient, making AI accessible around the world. Company Overview Cerebras is an AI company. We design processors for AI training and inference. We build AI systems to power, cool, and feed the processors data. We develop software to link these systems together into industry-leading supercomputers that are simple to use, even for the most complicated AI work, using familiar ML frameworks like PyTorch. Customers use our supercomputers to train industry-leading models. We use these supercomputers to run inference at speeds unobtainable on alternative commercial technologies. We deliver these AI capabilities to our customers on premise and via the cloud. AI compute is comprised of training and inference. For training, many of our customers have achieved over 10 times faster training time-to-solution compared to leading 8-way GPU systems of the same generation and have produced their own state-of-the-art models. For inference, we deliver over 10 times faster output generation speeds than GPU-based solutions from top CSPs, as benchmarked on leading open-source models. This enables real-time interactivity for AI applications and the development of smarter, more capable AI agents. The Cerebras solution requires less infrastructure, is simpler to use, and consumes less power than leading GPU architectures. It enables faster development and eliminates the complex distributed compute work required when using thousands of GPUs. Cerebras democratizes AI, enabling organizations that have less in-house AI or distributed computing expertise to leverage the full potential of AI. The rise of AI presents a unique set of compute challenges. Unlike other computational workloads, both training and inference require a huge number of relatively simple calculations, the results of which necessitate constant movement to and from memory, and to and from millions or tens of millions of compute cores. This traditionally demands hundreds or thousands of chips, and puts tremendous pressure on memory, memory bandwidth, and the communication fabric linking them all together. Cerebras started with a simple question: How can we design a processor, purpose-built to meet these exact challenges? If we were to start with a clean sheet, how would we avoid carrying forward the tradeoffs made for graphics and other workloads, and ensure that every transistor is optimized for the specific challenges presented by AI? Our answer is wafer-scale integration. Cerebras solved a problem that was open for the entire 75-year history of the computer industry: building chips the size of full silicon wafers. The third-generation Cerebras Wafer-Scale Engine (the “WSE-3”) is the largest chip ever sold. It is 57 times larger than the leading commercially available GPU. It has 52 times more compute cores, 880 times more on-chip memory (44 gigabytes), and 7,000 times more memory bandwidth (21 petabytes per second). The sheer size of the wafer-scale chip allows us to keep more work on-silicon and minimize the time-consuming, power-hungry movement of data. This enables Cerebras customers to solve problems in less time and using less power. Our AI compute platform combines processors, systems, software, and AI expert services, to deliver massive acceleration on even the largest, most capable AI models. It substantially reduces training times and inference latencies, while reducing programming complexity. Our business model is designed to meet the needs of our customers. Organizations seeking control over their data and AI compute infrastructure can purchase Cerebras AI supercomputers for on-premise deployment. Those that want the flexibility of a cloud-based platform can purchase Cerebras high-performance AI compute via a consumption-based model through the Cerebras Cloud, or via our partner’s cloud. We offer customers the flexibility to choose the solution that best aligns with their budgetary, security, and scalability requirements, and some customers choose to use both options simultaneously. 102


We have established a growing set of customer engagements spanning CSPs, leading enterprises, Sovereign AI programs, national laboratories, research institutions, and other innovators at the forefront of AI. While a substantial portion of our current business is supported by one primary customer, we are actively seeking to expand our reach and diversify our customer base. We collaborate with our customers to harness the power of AI to tackle their most significant challenges and drive breakthroughs across industries. Bloomberg Intelligence estimates that the AI market will grow to $1.3 trillion by 2032. Consumer and enterprise models like Google’s Gemini, Meta’s Llama, and OpenAI’s ChatGPT have driven demand for AI infrastructure training and inference solutions, powering AI applications such as specialized assistants, agents, and services. We believe that our AI compute platform addresses a large and growing AI hardware and software opportunity across training and inference, as well as software and expert services. We believe that further adoption of AI, accelerated by the advent of GenAI, and the widespread integration of AI into business processes, will rapidly expand our total addressable market (“TAM”) from an estimated $131 billion in 2024 to $453 billion by 2027, a compounded annual growth rate (“CAGR”) of 51%. We have experienced rapid growth, with revenue of $78.7 million and $24.6 million for the years ended December 31, 2023 and 2022, respectively, representing year-over-year growth of 220%. During the six months ended June 30, 2024 and 2023, we generated $136.4 million and $8.7 million in revenue, respectively. Since our inception, we have incurred operating losses and negative cash flows to develop, market, and expand our product portfolio and to continue our research and development activities. Our net loss for the years ended December 31, 2023 and 2022 was $127.2 million and $177.7 million, respectively, representing a year-over-year reduction of 28%. Our net loss for the six months ended June 30, 2024 and 2023 was $66.6 million and $77.8 million, respectively, representing a year-over-year reduction of 14%. Industry Background Over the past 40 years, the computer industry has followed a clear pattern: as major new computational workloads with distinct characteristics emerged, so too have new compute architectures. For example, the general-purpose needs of personal computing led to the x86 CPU. The low-energy needs of mobile devices resulted in the widespread adoption of ARM. Advancements in graphics demanded greater rendering parallelism, resulting in the creation of the GPU. With each new compute paradigm, technologists first attempted to adapt existing compute architectures to these workloads. But in each case, a new purpose-built architecture was ultimately needed to unlock the potential of the new paradigm. We believe this pattern is continuing with the rise of AI – the next major compute workload, with its own unique computational demands. In 2023, IDC estimated that the worldwide economic impact of GenAI would be close to $10 trillion by the end of 2033. This growth has been accelerated by the emergence of GenAI, a new class of powerful AI models that can create new content and reason across broad domains and multiple data types. These breakthrough capabilities translate to tremendous potential value creation and have driven rapid GenAI adoption. For example, ChatGPT took only five days to reach one million users (a feat that took Instagram 10 weeks and Twitter two years), and similarly, Meta’s open-source Llama 2 model attracted hundreds of thousands of AI developers within days of its release. The Computational Demands of Training and Inference The explosion of GenAI in the past 18 months has been likened to the “iPhone” moment, marking a pivotal shift in the industry. GenAI is already consuming compute resources at an unprecedented rate, faster than any other workload in history. As more powerful models are created and new use cases for GenAI are brought to market, we believe that demand for powerful and efficient AI compute solutions will continue to grow rapidly. The recent growth in the U.S. data center market supports this belief. Over the course of 2023, demand for U.S. third-party data centers has grown by nearly 50%, with AI driving significant growth. Major companies like Microsoft, Google, and Amazon are driving 500+ MW deals, whereas in 2015, data center deals were commonly in the 5 MW range. 103


Both training and inference demand immense compute, each with unique compute, memory, and memory bandwidth requirements. They represent two stages in the continuous lifecycle of AI models. Once trained, a model is “served” in production and used for inference. As part of this cycle, models in production are continuously being optimized to use fewer compute resources, and while that is happening, new and more powerful models are being trained, leading to the eventual obsolescence of the previous model—starting the cycle over. !business1a.jpg

*business1a.jpg*

Model Optimization Is Important, but As You Optimize, Models Continue to Evolve For training, the compute required is a function of a model’s size (number of parameters) and the amount of data used (number of tokens). The most capable GenAI models today have trillions of parameters and are trained on trillions of tokens. These models demonstrate superior accuracy and capabilities compared to small models due to their greater capacity to discern nuanced patterns from the data. As the industry has pushed to achieve greater AI capabilities, the size of GenAI models has grown, and we expect this trend to continue. The increase in model and data sizes has led to a dramatic surge in computational demand. As shown below, requirements to train GenAI models have grown 40,000-fold in just five years. Today, training for GenAI requires enormous GPU clusters, sophisticated engineering teams, and months of time for a single run. A training run can cost over a hundred million dollars, and improving the model and keeping it fresh with new data requires additional fine-tuning and regular re-training. 104


!business2a.jpg

*business2a.jpg*

Illustration of Growing Compute Requirements Over Time (ExaFLOPS to Train LLMs) For inference, the required compute is a function of model size, user demand, and time spent on inference. Larger models use more computational resources, and each user request also increases the compute need. Dedicating more compute resources to inference also generally results in better output quality, as more compute enables deeper chains of thought and reasoning. These factors contribute to significant and ongoing operational costs. We expect the demand for inference to grow, especially as larger and more capable models become more widely adopted. There is currently a direct tradeoff between model capability and responsiveness of user experience because the largest and most capable models demand more inference compute and therefore run more slowly during inference. As inference speeds improve, larger GenAI models can be deployed in a responsive manner, expanding their use in real-time applications and agent-based systems. Faster inference can enable multiple inference requests to be made to the same or different GenAI models, with each request building upon the results of the previous request, all within the time it previously took to execute a single inference request. Advancements in model reasoning capabilities are equally important. For instance, OpenAI’s o1 model, released in September 2024, uses chain of thought—a reasoning process that leverages multiple intermediate inference steps—to solve complex coding, math, and science problems that were unsolvable by earlier models. We expect that increasing compute resources used for inference by orders of magnitude to enable deeper chains of thought will continue to yield significant improvements in output quality for problems that require complex reasoning. These reasoning improvements enable models to tackle more sophisticated tasks, broadening their range of use cases. We believe that as both inference speed and reasoning capabilities advance, GenAI models will support more demanding applications, thereby expanding the market opportunity for inference. As of June 2024, we estimate that 40% of the AI data center market is attributable to inference, and we expect this to grow as more AI applications and products are brought to market. Existing Compute Architectures Are Fundamentally Limited for GenAI Training and Inference GPUs, though better than CPUs for AI workloads, face fundamental limitations when processing the unique characteristics of large GenAI models. In graphics, the calculation for each pixel is often independent of every other pixel. This means that for large graphics workloads, many GPUs can easily execute on separate parts of the rendering problem, without a need for high levels of interdependent communication. The AI workload is different. GenAI models are complex, interconnected compute graphs that require the constant communication of intermediate calculations to train. This requires a high amount of data movement to and from memory and across cores. Similarly, during inference, these models generate outputs sequentially – each 105


dependent on the previous output – requiring the full model to be constantly moved in and out of memory to produce successive outputs, and again requiring massive data movement between cores and memory. GPUs face inherent scalability and complexity challenges when faced with the distinct, communication-heavy requirements of GenAI workloads. For Training – Individual GPUs Are Too Small, and Scaling to Many GPUs is Highly Inefficient Large GenAI models far exceed the memory and processing limits of a single GPU. For example, to train GPT-3, which OpenAI has said required ~3.14 x 10^23 floating point operations, it would take a single NVIDIA H100 more than eight years of running at peak theoretical performance to train the model. Recent models like GPT-4 and Gemini are over 10 times larger in parameter size than GPT-3. Consequently, training a large GenAI model on GPUs in a tractable amount of time requires breaking up the model and calculations, and distributing the pieces across hundreds or thousands of GPUs. These hundreds or thousands of GPUs then need to constantly communicate with each other across a network, creating extreme communication bottlenecks and power inefficiencies. This distributed compute problem also creates a high level of complexity for developers, who are responsible for partitioning and coordinating the compute, memory, and communication across GPUs, so that they can work together in a complex choreography. This is an ongoing cost and slows down time-to-solution, as the delicate balance of bottlenecks needs to be reconfigured every time the ML developer wants to change the model architecture, model size, or run on a different number of GPUs. For many organizations, distributed programming is one of the highest barriers to entry for leveraging GenAI. !business3a.jpg

*business3a.jpg*

Large GenAI Models Must Be Divided and Coordinated Across Thousands of GPUs, Leading to Communication Bottlenecks and Developer Complexity For Inference – GPU Efficiency is Low and Limited by Memory Bandwidth During generative inference, the full model must be run for each word that is generated. Since large models exceed on-chip GPU memory, this requires frequent data movement to and from off-chip memory. GPUs have 106


relatively low memory bandwidth, meaning that the rate at which they can move information from off-chip HBM to on-chip SRAM, where the computation is done, is severely limited. This leads to low performance as GPUs cores are idle while waiting for data – they can run at less than 5% utilization on interactive generative inference tasks. Low utilization and limited memory bandwidth impact the responsiveness and throughput of GPU-based systems and hinders real-time applications for larger models. This can limit the capability and adoption of emerging inference applications which are especially latency-sensitive, such as code generation and multi-turn AI agents that need to string together multiple calls to different GenAI models. This inefficiency also necessitates larger GPU deployments and dramatically drives up the cost of inference. GPU companies have tried to address these challenges, but the issues of small core count, memory size, and memory bandwidth are fundamental hardware limitations that persist. Interconnect technologies like InfiniBand, PCIe, and NVLink are limited by their physical interfaces, and moving data across them is thousands of times slower and more power-hungry than keeping and moving the data on silicon. Software libraries intended to simplify distributed computing still require developers to manage complex parallelism strategies and extensive codebases. Realizing the physical challenges of repurposing small chips for a large compute problem, the GPU industry has recently announced new multi-chip packaging techniques, but these also yield only minimal expansions of GPU chip size. Accelerating GenAI requires a dedicated compute solution, designed for the unique requirements of GenAI, that can deliver faster training times, real-time inference speeds, and simple developer workflows, at reasonable cost. Our Solution We believe Cerebras has built the world’s fastest commercially available AI training and inference solution. Our dedicated AI hardware and software platform is powered by the Cerebras Wafer-Scale Engine – a processor the size of an entire silicon wafer that brings more on-chip compute, memory, and bandwidth resources together than any other commercially available processor in the semiconductor industry. A single WSE replaces a cluster of GPUs, reducing the time-consuming, power-hungry movement of data, removing the need for complex distributed programming, and providing exceptional compute speed. Compared to the leading 8-way GPU system of the same generation, many of our customers have achieved over 10 times faster training time-to-solution. Our inference offering is over 10 times faster than GPU-based solutions from top CSPs, as benchmarked on leading open-source models. A single CS-3 system delivers three times more compute per watt than the leading 8-way GPU system. Our solution consists of the following elements: Cerebras Wafer-Scale Engine (WSE). At the heart of our solution is the Cerebras Wafer-Scale Engine, the largest chip ever sold. Our third generation WSE, the WSE-3, is 57 times larger than the leading commercially available GPU (NVIDIA H100) and has 52 times more compute cores, totaling 900,000. It features 880 times more on-chip memory (44 gigabytes SRAM) and 7,000 times more memory bandwidth (21 petabytes per second) than the leading commercially available GPU. Developing the WSE required overcoming decades-long-standing industry challenges, including inter-die connectivity, yield optimization, efficient packaging, power management, and advanced cooling. 107


!business4ba.jpg

*business4ba.jpg*

Cerebras Wafer-Scale Engine 3 Versus NVIDIA H100 Size Comparison Image is illustrative; scale is approximate. !business5a.jpg

*business5a.jpg*

Cerebras Wafer-Scale Engine 3 Versus NVIDIA H100 Capability Advantage The immense scale of the WSE delivers significant acceleration, efficiency, and simplicity for AI training and inference. 108


For Training. Each Cerebras WSE has enough compute and on-chip memory to run even the largest, multi-trillion parameter GenAI models on a single chip, thus avoiding the complexities of chip-to-chip data movement. This is unlike GPUs, which require users to fragment large models across multiple processors and deal with complex inter-GPU communication. To further speed up training time-to-solution, Cerebras users can simply add more WSEs to the problem. Since each WSE can fit the whole model, multiple WSEs can accelerate model training simply by having each chip independently work on a subset of the training data. Because the model is never split across the WSEs, no complex model distribution or carefully orchestrated communication is needed. The elegance of this architecture is designed to allow users to effortlessly increase training speeds with near-linear performance scaling as more WSEs are added. For Inference. Wafer-scale integration keeps all critical data on-chip and close to compute cores, resulting in 7,000 times more memory bandwidth than the leading GPU solution. This allows the WSE-3 to deliver over 10 times lower latency for real-time GenAI inference, which means an inference response that takes 10 seconds to generate on GPU-based platforms from top CSPs takes only one second to complete using the Cerebras solution. The WSE-3 can accomplish this speedup at vastly lower power consumption, as retrieving one bit of data from on-chip SRAM on 5nm silicon requires only 1% the energy that is needed to do the same from off-chip HBM. !business6ba.jpg

*business6ba.jpg*

Cerebras Inference Is the Fastest Inference Solution on Llama 3.1 8B 109


!business7ba.jpg

*business7ba.jpg*

Cerebras Inference Is the Fastest Inference Solution on Llama 3.1 70B Cerebras System (CS). The CS AI computer system houses the WSE and delivers innovative power and cooling to the chip. Our third generation CS (“CS-3”) delivers three times more compute per unit power than the leading 8-way GPU system (NVIDIA DGX H100). This compact AI powerhouse is designed to easily integrate into standard data centers – it occupies less than half a standard rack (16RU/27 inches tall), and connects into the network via standards-based 100G Ethernet. !business8a.jpg

*business8a.jpg*

The CS-3 System, with Innovative Power and Cooling, is Designed to Fit in a Standard Data Center Rack 110


Cerebras AI Supercomputer. The Cerebras AI Supercomputer is designed to streamline scaling up to 2,048 CS-3 systems for maximum AI acceleration, with more efficiency and simplicity than scaling up to large GPU clusters. Unlike with GPUs, where users must break up and distribute their model across many chips, thereby introducing the need to communicate across those compute elements, each WSE can run the model in its entirety. In turn, a Cerebras AI Supercomputer only needs to spin up another copy of the full model on each additional CS system to process the training dataset more quickly. This enables a near-linear performance increase as CS systems are added to a problem, takes only seconds to configure, and does not incur the overhead or complexity of heavy inter-chip, inter-system communication. Our scalable execution model is designed to simplify the development workflow for large GenAI training. We designed our AI Supercomputer to enable users to elastically scale workload computing resources up to 256 exaFLOPS (2,048 CS-3 systems) just by changing a single number in their code, allowing users to program as if for a single powerful device. Training a GPT-3 sized model on Cerebras, for example, uses 97% fewer lines of code compared to on clusters of GPUs, greatly accelerating the speed of AI model developments for larger-scale models. The combined power and scaling simplicity of the Cerebras AI Supercomputer provides industry-leading training and inference speeds for even the largest and most complex GenAI models, while obviating the need to invest thousands of programmer hours into distributed programming work. !business9a.jpg

*business9a.jpg*

Zero-Complexity Scaling on Cerebras Versus GPUs: Training a GPT-3 Sized Model With Cerebras Requires 97% Fewer Lines of Code Compared to on Clusters of GPUs 111


!business10a.jpg

*business10a.jpg*

A Cerebras Supercomputer Deployment at One of Our Colocation Facilities Cerebras Software Platform (CSoft). Our proprietary software platform, CSoft, is core to our solution and provides intuitive usability and improved developer productivity. CSoft seamlessly integrates with industry-standard ML frameworks like PyTorch and with popular software development tools, allowing developers to easily migrate to the Cerebras platform. CSoft automatically saves model outputs in a standard “checkpoint” format, enabling users to resume training or inference on model work started on other hardware platforms. This checkpoint compatibility extends to open-source models available in the popular HuggingFace repository. CSoft eliminates the need for low-level programming in CUDA, or other hardware-specific languages. Starting from a user’s PyTorch model, the CSoft graph compiler automatically maps model operations to the WSE, creating an optimized executable without user-level intervention. CSoft is co-designed with the wafer-scale hardware architecture to provide programming efficiency and simplicity. It allows ML users to accelerate training and inference on models of any size, scaled across any configuration of the Cerebras AI Supercomputer, just by changing one number in a configuration file, simulating a single-device programming experience without the complexities of distributed programming. This drastically reduces operational overhead and speeds up developer iteration time and business impact. Cerebras Inference Serving Stack. The Cerebras Inference Stack is built on top of CSoft and is designed to allow customers to easily deploy even the largest GenAI models at industry-leading inference speeds. With the Cerebras Inference API, developers are able to easily point their applications to popular or custom models, just by changing their API key. This simple API, which closely matches the interface design of the popular OpenAI API, is designed to facilitate rapid developer adoption and ease of use. We believe there will be stickiness to the Cerebras platform as customers experience industry-leading inference performance for their interactive applications. Our serving software automatically handles system-level optimizations for our inference solution and is designed to enable low latency and high cost effectiveness. 112


AI Model Services. Our AI model services further amplify speed to solution. Our team of AI practitioners helps customers design research experiments, train models, and optimize processes designed to achieve the fastest time-to-solution. These services complement our advanced hardware and software platform, providing an end-to-end solution for rapid and efficient AI development and deployment. We believe we are one of a select number of companies in the world that has trained high-quality, large GenAI models on massively parallel compute clusters. We have contributed state-of-the-art models into the open-source community and have published widely on AI methodology and practice. Our team’s work spans across model architectures, parameter sizes, and data types – including text, time series data, code, biological and other sequence data, electronic healthcare records, medical imagery, radio frequency, and radar data. Cerebras AI researchers have trained LLMs in English, Arabic, Spanish, Japanese, and Catalan. We excel at helping customers translate AI potential to business impact. Our AI experts guide customers in developing custom AI strategies, designing and building AI models, and applying cutting edge AI techniques to achieve high-quality results, thereby delivering downstream business impact. We augment our customers’ existing teams with specialized AI capabilities, and the models we build together routinely beat the existing state of the art, providing customers with a distinct competitive advantage. Summary of key customer benefits include: Enables over 10 times faster training time-to-solution compared to leading 8-way GPU systems of the same generation, as reported by many of our customers. This dramatically accelerates AI model time-to-solution, enabling businesses to test ideas faster, iterate more quickly, and bring next-generation GenAI-powered products and services to market, faster and more cost effectively. Delivers over 10 times faster GenAI inference compared to GPU-based solutions from top CSPs, as benchmarked on leading open-source models. The WSE’s massive on-chip SRAM capacity keeps the vast majority of memory-to-compute communication on-silicon, thereby avoiding the memory bandwidth bottleneck faced by GPU-based solutions. Our resulting ultra-low latency delivers industry-leading inference speeds and real-time responses back to users, even on large, cutting-edge GenAI models. Ten times more speed compared to GPU-based solutions from top CSPs also allows developers to make ten times more inference calls in the same amount of time. This supports a new level of model capability delivered by techniques like multi-step inference and agentic AI flows, which leverage more inference calls to produce higher reasoning capability for more complex tasks in domains such as coding, math, and science applications. End-to-end solution. Cerebras offers a unified platform purpose-built to accelerate fundamental compute characteristics of both AI training and inference. While many other emerging AI chips companies have chosen to specialize in only one phase to allow for optimization tradeoffs across the axes of compute, memory, memory bandwidth, and simplicity of use, Cerebras excels across these key dimensions, made possible by wafer-scale integration. This allows customers to swiftly transition from training to fine-tuning to deploying high-quality GenAI models on the same platform, eliminating the need for investing in and maintaining separate hardware infrastructure. Zero distribution complexity. The biggest challenge in training on GPU clusters is the complexity of distributed programming across multiple devices and systems. This often requires large engineering teams weeks or months of optimization, and results in limited performance scaling. Using Cerebras, users can effortlessly run a GenAI model of any size. It takes no additional code to achieve automatic near-linear performance scaling across the nodes of a Cerebras Supercomputer, unlike the 20,000+ lines of distributed programming code required to scale large models across similar memory and compute resources on large GPU clusters. Low migration cost. Our proprietary CSoft platform integrates seamlessly with familiar ML frameworks and tools, like PyTorch, eliminating the need for AI teams to learn new languages or adapt to new development environments and easing the transition from other hardware platforms. Power and operational efficiency. Cerebras outperforms GPU systems in power efficiency due to both hardware and architectural advantages. Wafer-scale integration allows the majority of AI workload communication 113


to remain on-chip, significantly reducing data movement distance and power consumption; moving a bit of data on the WSE-3 takes less than 1% of the energy needed to do the same over off-chip GPU interconnects. This drives significant operational cost savings (CS-3 draws one-third of the power of the leading 8-way GPU system) and streamlined management for organizations deploying AI at scale. Expert-led model training and AI integration services. We offer expert-led foundation model training, fine-tuning, and retrieval-augmented generation services to customers. Our team provides guidance on cutting-edge AI methods that work on top of our hardware, assisting customers to derive the maximum value from their AI investment. Examples of Customer and Industry Impact We have an expanding customer base that includes leading enterprises, Sovereign AI initiatives, cloud service providers, government agencies, and research institutions at the forefront of AI and at the intersection of AI and HPC. These customers leverage Cerebras to tackle complex AI challenges and achieve previously unattainable business outcomes, even with the most advanced GPU systems. Our customers find fundamental value in the simplicity of the Cerebras solution. It unlocks breakthrough business and scientific use cases by removing limitations in development time, programming complexity, and runtime speed. For example, a leading pharmaceutical company trained an epigenomic language model using Cerebras and noted that the training speedup afforded by the Cerebras system enabled them to explore architecture variations, tokenization schemes and hyperparameter settings in a way that would have been prohibitively time and resource intensive on a typical GPU cluster. As we continue to expand our go-to-market capabilities and more customers recognize the benefits of our solution we expect our pipeline to grow. The tangible benefits of our products are evident in the strong gains in performance and time-to-value experienced by our customers across various industries: A global AI technology group (G42) reduced AI model convergence time for a 13 billion parameter Arabic language model from 68 days on a GPU cluster to just four days using Cerebras, improving training time by 17 times while also achieving higher model quality. They furthered this work by subsequently training a state-of-the-art bilingual Arabic-English 30 billion parameter model on 1.6 trillion tokens, now being served on the Microsoft Azure AI Model Catalog. A leading pharmaceutical company reduced training time on an epigenomic language model from 24 days to just 2.5 days using Cerebras, achieving a 90% reduction in time-to-solution. This resulted in a new model that was the first of its kind to combine both DNA data and epigenetic state data of chromosomes for 127 different cell types. This model exceeded the state-of-the-art on multiple benchmarks and is being used by the customer to facilitate further study of gene regulation and function. A major U.S. healthcare provider leveraged Cerebras AI services and compute to develop and train an industry-leading medical model in eight weeks, beating open-source medical model benchmarks, and upskilling their in-house AI team through partnership with our AI expertise. A U.S. national lab trained a genome-scale foundation model, predicting the evolution of emergent variants of SARS-CoV-2 virus, 19 times more quickly on a single Cerebras system compared to the leading 8-GPU system of the time (152 times faster than a single GPU). The same national lab trained an AI model using Cerebras that could not converge on any number of GPUs. This model won the 2022 Gordon Bell Special Prize for High Performance Computing-Based COVID-19 Research. Another U.S. national lab showed 179 times faster performance on a molecular dynamics simulation using a single Cerebras system, compared to the entirety of Oak Ridge National Laboratory’s Frontier, the largest 114


supercomputer in the world, with more than 37,000 GPUs and reportedly costing approximately $600 million to build. A U.S. government agency is leveraging our real-time inference capabilities to create the world’s first, large-scale, virtual radio frequency simulation environment for developing, training, and testing advanced radio frequency systems. Our Business Model We use a combination of direct sales and partnerships to address the rapidly expanding AI market. We offer both on-premise solutions and cloud-based solutions to provide maximum flexibility to our customers. We offer a collection of services from data center deployment to AI expert services and AI Supercomputer operation and management, to provide our customers with the support they need to train, deploy, and accelerate GenAI time-to-value. On Premise. Our direct-to-customer sales force sells our AI Supercomputers to leading organizations who seek maximum control over their data and their AI infrastructure, fulfilling their needs for high-performance AI compute on premise. For on-premise use, we offer deployment services, as well as a subscription to an ongoing stream of software updates and upgrades. Our on-premise AI Supercomputers support both training and inference. They can be configured either for both workloads, or to be further optimized for only one, depending on our customers’ needs. Cloud-Based Computing Services. We sell Cerebras solutions via our cloud offering as well as via the Condor Galaxy Cloud owned by Group 42 Holding Ltd (together with its affiliates, “G42”), our partner CSP. Our cloud solutions provide customers fast and flexible access to our powerful AI acceleration hardware, with payment based on time or work (number of models). This offering gives our customers the ability to train LLMs with extraordinary speed and deploy them for inference at ultra-low latencies, all without the complexity or time needed to build and manage on-premise infrastructure. Cerebras Inference Cloud. Our real-time inference solution is also available via a dedicated inference cloud service. Leveraging our Cerebras Inference Serving Stack, this cloud API offering allows developers to directly point their applications towards efficient and reliable model serving endpoints. On Cerebras Inference Cloud, we host both popular open-source models and proprietary customer models. For customers who do not need direct compute access and are not interested in managing their own inference serving software stack, our inference cloud offering is the quickest and simplest way to leverage our fast model inference services. We provide a combination of these offerings to customers who may benefit from leveraging both on-premise and cloud-based options—for example, enabling them to quickly use the cloud for their largest training jobs, while enjoying the cost efficiencies of on-premise infrastructure for their baseline AI work. This flexibility allows customers to choose the solution that best aligns with their budgetary, security, and scalability requirements. Additionally, customers can train models on-premise and then leverage our inference cloud for production, benefiting from flexible serving resources that can adapt to fluctuating demand. This end-to-end solution allows seamless integration from training to production, serving the entire lifecycle of a customer’s AI needs. We also provide professional services to assist customers throughout the AI workflow. From developing strategy, designing and building models, to deploying and maintaining the final models either on-premise or in the cloud, we help customers achieve optimal outcomes. Our Market Opportunity We participate in a large and growing AI market. Enterprises, research organizations, and governments alike are developing AI initiatives to rapidly evolve and drive efficiencies across their entities. As estimated by McKinsey in 2023, GenAI use cases may add up to $4.4 trillion to the global economy on an annual basis. Our full suite of AI computing solutions addresses use cases for training, inference, software, and expert services. We estimate the TAM 115


for our AI computing platform to be approximately $131 billion in 2024, growing to $453 billion in 2027, a 51% CAGR. This TAM is comprised of the following core markets: AI Compute for Training. Demand for GenAI solutions has grown significantly across virtually all use cases. In a recent Gartner, Inc. poll of more than 1,400 executive leaders, 45% reported that they are in piloting mode with GenAI, and another 10% have put GenAI solutions into production (Source: Gartner®). Significant capital is being dedicated to evaluating GenAI solutions and the impact they could have on everything from productivity, efficiency, to customer support. As businesses continue to evaluate and deploy solutions, we believe the market for training new models will continue to grow. Based on market estimates in Bloomberg Intelligence research, our estimate of the TAM for AI Training Infrastructure is $72 billion in 2024, growing to $192 billion in 2027, a 39% CAGR. This segment of the market includes AI Infrastructure-as-a-Service. AI Compute for Inference. While GenAI training is essential to developing models that are powerful and accurate, we believe inference is the next phase of the ongoing wave of AI disruption. As more enterprises develop models and start to deploy their models in applications at scale, the need for high performance and efficient inference is becoming critical to fully realize the commercial potential of ML. The demand for inference compute is driven by the number of users and frequency that inference is used. We believe that the inference opportunity is enormous, as the market is in the early phases in its adoption cycle. Our estimate of the TAM for AI Inference is $43 billion in 2024, growing at an estimated 63% CAGR to $186 billion in 2027. Software and Expert Services. Based on market estimates in Bloomberg Intelligence research, our estimate of the TAM for GenAI Software and Services is $16 billion in 2024, growing to $75 billion in 2027, a 67% CAGR. We believe we are at the very early stages of a large and fast-growing market opportunity. As adoption of AI continues to accelerate, we expect numerous new applications will be identified, and we believe our solutions are well-positioned to capitalize on the wave of disruption that will come in the coming years. Customers and Ecosystem Our customers include leading enterprises, Sovereign AI initiatives, cloud service providers, government agencies, and research institutions at the forefront of AI and at the intersection of AI and HPC. The majority of our revenue for 2023 is from the sale and deployment of our AI systems, but we expect significant traction in onboarding new customers onto our cloud offering. The scalability and flexibility of our solutions enables customers to leverage both hardware and cloud solutions concurrently, allowing them to adapt their approach as their needs grow. Our customers also leverage our AI model services to achieve state-of-the-art AI results. We sell our products directly to our customers and in some cases, we rely on fulfillment partners for deployment and logistical purposes. Additionally, we share our innovations and research on open-source platforms like HuggingFace, where our published models and datasets have garnered significant traction. An example is our open-source BTLM Model, which was downloaded over one million times. This broad engagement keeps us at the cutting edge of AI technology and fosters a collaborative relationship with the ML community. Competitive Strengths We engineer our purpose-built AI systems to be faster and more capable than any other solution available in the market today. We believe our design is capable of meeting today’s needs and is scalable to address tomorrow’s challenges. Our competitive strengths include: The world’s first and only wafer-scale chip in the market. Achieving wafer-scale success has been a goal of the semiconductor industry for decades, and we are the first company to overcome and solve the fundamental technical challenges required to productize a wafer-scale chip, including inter-die connectivity, yield, packaging, power, and cooling. Our wafer-scale chip architecture combines a massive amount of fast, on-chip memory, directly 116


next to vast compute resources on the same piece of silicon. It is capable of running the largest GenAI models on a single wafer-scale chip instead of tens, hundreds, or thousands of competitor devices. This eliminates the need for distributed computing while running large GenAI models, enabling AI developers to use up to 97% less code when working with large models on our platform compared to on clusters of GPUs and greatly accelerating the speed of AI model development for larger-scale models. The details of the hardware are completely abstracted away from the AI developer, allowing them to focus on model development and deployment, which they are empowered to do faster than on any other system. Full system solution that is easy to deploy and efficient to operate. Built with our wafer-scale chip at its core, we design full systems that seamlessly fit into data center racks. Unlike many other players in this industry, we both design the chip and the system around it. By designing the processor and the system together, we are able to address thermal and power delivery challenges by optimizing the full system with our proprietary power delivery and cooling technology. This allows us to keep the system operating efficiently, optimizing the energy consumption and underlying operating costs for our customers. Comprehensive software suite that leads to ease of adoption and shortens time to deployment. Developed over eight years by our world-class team, the Cerebras Software platform allows for seamless programmability of AI models through our integration with PyTorch, the industry-standard ML framework for AI developers. Our tools allow users to bring models that have been trained on other hardware onto our platform for training or for inference. Likewise, users can train models on our platform and deploy those models for inference elsewhere. This increases ease of adoption and shortens time-to-deployment. Our AI platform addresses both training and inferencing markets. We are strongly positioned to provide an AI platform that is powerful, efficient and easier to deploy for both training and inferencing workloads. For training, our AI platform enables customers to effortlessly and swiftly use the most advanced GenAI models available on the market. It allows any single user to effectively develop and refine models with multi-billion-to-trillion parameters, without the need for specialized software frameworks or help from distributed computing experts. For inferencing, given our advantages in memory capacity, our AI platform delivers industry-lowest latencies and high generation throughput. This helps our customers to unlock cutting edge performance for ultra-low-latency use cases leveraging GenAI, such as real-time, interactive user applications and intelligent AI agents that often require processing data across multiple GenAI models and APIs. Scalable architecture. We have developed our solutions to support models and data sets of large and varying sizes. Our current generation CS-3 is designed to support models with up to 24 trillion parameters, much larger than even today’s state-of-the-art GenAI models. Our platform is designed to seamlessly scale from 1 to 2,048 CS-3 systems, forming an AI supercomputer that is even further differentiated by our proprietary interconnect and memory technology. This enables customers to seamlessly increase compute resources from small-scale experiments to large-scale deployments. Our future-ready design enables us to support the most demanding AI applications and are designed to ensure that our customers’ investments remain forward-compatible as new advancements in AI emerge. We expect this benefit to be crucial for long-term customer satisfaction and retention. AI model services help customers translate AI potential to business impact. We provide customers with AI model services to help them develop solutions that are customized to meet their needs and help them realize the full value of their AI investments. These services include model selection, data preparation, training, and solutions integration. World-class AI talent with a proven track record of innovation and execution. We have assembled a world-class team of industry leaders in integrated circuit design, processor architecture, power delivery, cooling, system engineering and software. Over the last five years, we have introduced three generations of our WSE, each time achieving two times the performance of its predecessor, and bringing new IC, power, and cooling technologies to bear. This track record of innovation and execution demonstrates our singular focus on solving the critical performance challenges of AI compute, and systematically improving our product to achieve more performant AI compute. Our research and development organization was over 80% of our total headcount as of December 31, 2023, with approximately three-quarters of research and development headcount consisting of software engineers. 117


Growth Strategies We believe we are positioned for sustained growth in the rapidly expanding market for AI acceleration solutions. We have designed our focused strategies to drive continued success and establish ourselves as a long-term leader: Increase sales to our existing customers. We have established a strong land-and-expand track record with our existing customer base, which comprises leading enterprises and research institutions looking to harness the potential of AI. We intend to deepen these relationships by expanding our product and service offerings tailored to their evolving needs. Our strategy focuses on demonstrating the value proposition of our solutions through initial engagements, cross-selling complementary products, and identifying new use cases within existing customers. Our deep technical expertise and customer-focused professional services consistently deliver value beyond initial expectations. This success helps drive increased investment from existing clients. Expand our customer base. We believe a core component of our growth strategy centers on expanding our diverse customer base across industries and applications. We plan to aggressively pursue opportunities in relevant sectors such as healthcare, pharmaceutical, biotechnology, government, financial services, sovereign, and energy, to name a few, where our AI acceleration capabilities can address critical computational bottlenecks. We will seek to drive this expansion by focused sales and marketing initiatives, highlighting the transformative potential of our technology with targeted use cases. We intend to leverage our existing success stories and strategic partnerships to both bolster our credibility within new markets and establish key channels for customer acquisition. For example, we entered into a partnership with a world-leading European AI company to create sovereign and secure AI solutions for governments and enterprises resulting in a first-time buy of over $4 million. Additionally, we will continue investing in the development of tools and resources that streamline the onboarding experience for new customers to enable seamless adoption of our solution. Further penetrate into the rapidly growing inference market. We recently launched our inference solution. A large opportunity for our growth is to make our inference solution widely available. The immense amount of memory bandwidth and capacity on our chip allows us to deliver significantly lower generation latency and higher generation throughput over GPUs. By making API-based inferencing available through our cloud offering, we could significantly accelerate our adoption into inferencing use cases. Based on the inference TAM of $43 billion in 2024, growing at an estimated 63% CAGR to $186 billion in 2027, we believe our expansion into inference will be a significant growth driver. Benefit from opportunities in large adjacent AI and compute-intensive markets. We intend to leverage our differentiated solution to address ever-evolving computing challenges in emerging end markets and applications. We are actively enabling applications in fields like life sciences, materials science, and financial modeling, where our cutting-edge AI computing solutions can unlock new discoveries and solve complex problems. Geographically, our strategy includes deepening our investment in partnerships with large Sovereign AI initiatives and new markets to accelerate AI adoption. For instance, AI’s economic value in the Middle East’s Gulf Cooperation Council countries is projected to be as much as $150 billion, or approximately 9% of such countries’ combined gross domestic product, according to a McKinsey report issued in 2023. We believe our strategic partnership with G42 positions our solutions at the forefront of innovation and expansion in these emerging AI markets. Accelerate our existing product roadmap as well as develop new products for emerging use cases. We believe continued technological innovation is a cornerstone of our future growth. Our substantial investment in research and development fuels advancements to further differentiate our wafer-scale technology and develop our software to optimize performance, accessibility, and ease of use of our platform. Our close partnerships with leading AI researchers, industry pioneers, and the open-source community grant us early insights into cutting-edge market developments. Moreover, as AI infrastructure requirements scale, we expect emerging use cases to require new products with added functionalities to solve data, networking, and memory bottlenecks. With our continued focus on innovation, we intend to develop and introduce new products and form factors that will enable us to service a larger portion of our market opportunity. 118


Advance product adoption by proliferating cloud deployment of our AI solutions. We intend to accelerate our growth by expanding access to our revolutionary AI systems through cloud deployment. We expect this strategy to make our unique, high-performance computing capabilities available to a significantly broader customer base. Cloud solutions reduce the upfront capital investment required for customers, enabling more rapid experimentation and wider adoption of our technology. By offering our technology as a cloud service, we can streamline workflows for data scientists and machine learning engineers, allowing them to focus on innovation instead of infrastructure management. Our Technology We believe we have built the fastest commercially available AI training and inference solution in the world. Our solution combines our industry-leading wafer-scale engine with a supercomputer system architecture and a co-designed software suite. Every component of our platform is specifically designed to meet the computational needs of AI workloads. We combine massive computational resources, fast on-chip memory, and high-bandwidth interconnect to simplify and accelerate AI training and inference. Our product offerings are available through a unified platform that is both easy to use and simple to deploy. Our platform’s exceptional performance is driven by a portfolio of fundamental technological innovations, that resolve challenges in the computing industry that had been open for decades. WSE-3 – The World’s Largest, Most Powerful Commercially Available Chip Throughout the history of computing, the drive towards larger chip sizes with more integrated components has consistently produced significant performance and efficiency improvements. This has driven the entire semiconductor industry towards Moore’s Law, which predicts that the number of transistors in an integrated circuit will double every two years. However, the unprecedented demand for AI computing has underscored the need for improvements that go beyond the traditional cadence of compute advancements. To meet this demand, we have designed the Cerebras Wafer-Scale Engine – the largest chip ever sold, and which has broken Moore’s Law. With four trillion transistors, the WSE-3 achieves a milestone that Moore’s Law did not predict would occur until 2034. This leap marks a significant advancement in chip design and manufacturing, setting new standards for computational power and efficiency. 119


!business11a.jpg

*business11a.jpg*

“Moore’s Law” Over the Past 40 Years The white line shows the Moore’s Law trend of the largest processor chips in the industry. The orange line shows the new trajectory created by the Cerebras wafer-scale technology. Wafer-scale integration is a unique capability. In the entire history of the computer industry, only Cerebras has delivered a wafer-scale product to market. To achieve wafer-scale integration, we invented and productized key chip design technologies: •Multi-die interconnect. Traditionally, small die—regions of silicon containing an integrated circuit—are replicated independently on a silicon wafer and then cut up (“diced”) into small, separate chips. We have developed the technology to connect these otherwise independent die together at the wafer level, at the semiconductor fabrication plant. The inter-die connectivity uses a special cross-reticle connection that is integrated into the overall fabrication process. •Fault-tolerant architecture. A primary factor in the commercial viability of a semiconductor is the yield. The fabrication process introduces tiny defects – and chips with these defects are discarded and considered yield loss. Recognizing that large chips are more prone to defects than small chips, we invented new techniques designed to withstand defects, rather than seek to avoid them. Combining architectural innovations with insights from adjacent industries, we used a combination of redundant compute cores and redundant routing to address the yield problems. The wafer behaves like a hyper-scale data center, using a “fail in place” mechanism to handle the inevitable failures at scale. Flaws are designed to be recognized, shut down, and routed around. Redundant cores are used to re-form a logically functional whole. The enormous scale of our wafer-scale engine provides four fundamental technical advantages for AI: 1.Massive, tightly-coupled compute. All modern, high-quality GenAI models require too much compute to run on a single GPU. Our wafer-scale processor is purpose-built for these AI workloads and tightly couples a massive amount of compute resources (900,000 cores, 52 times more than the leading GPU currently in the market) onto the same piece of silicon. By integrating the equivalent compute of a cluster of GPUs onto a single chip, the WSE can train or run inference on large AI workloads without breaking up the problem across a large number of small devices. Keeping so much compute in one place also allows the WSE to 120


achieve significantly higher performance and power efficiency, compared to GPU clusters that require heavy communication across hundreds or thousands of small chips. Put simply, the WSE avoids the bandwidth, latency, and power tax of distributing communication-heavy AI computation across chips. !business12b.jpg

*business12b.jpg*

Comparison of Cerebras Wafer-Scale Integration Versus Traditional GPU Cluster Integration 2.Fast, efficient interconnect. For both inference and training, AI requires extraordinary amounts of data movement. On GPU clusters, most of the data movement occurs between individual GPU chips using traditional I/O interconnects such as NVLink or Infiniband. On the Cerebras wafer-scale chip, most of the data movement is performed locally using the on-chip fabric that achieves nearly 30,000 times higher interconnect bandwidth and is over 100 times more power-efficient per bit of data moved, compared to leading GPU interconnects. This architecture achieves industry-leading interconnect performance and efficiency compared to current GPU servers. This directly translates into faster acceleration for both training and inference, as it minimizes communication across slow, physical networking interfaces. 3.Outsized on-chip memory and memory bandwidth. AI workloads use memory to store the model state (parameters) and active data samples running through the model (activations). Fast access to this data is critical to the performance of both training and inference. The WSE uses a unique memory-in-compute architecture, which integrates memory and processing cores on the same chip, minimizing the distance that data must travel between memory and compute. Memory-in-compute architectures provide a step-change improvement in bandwidth compared to traditional external-memory architectures. This directly results in dramatically reduced latency and power consumption. Traditionally, memory-in-compute architectures were restricted to only small-scale problems because the tight integration of memory and compute was constrained by the small size of single chips. With wafer-scale integration, we are able to scale to enormous on-chip memory (SRAM) capacity (44 GB) on a single chip, while achieving 7,000 times more memory bandwidth (21 petabytes per second) compared to the leading GPU. 121


!business13a.jpg

*business13a.jpg*

WSE-3 Achieves 7,000 Times More Memory Bandwidth Compared to NVIDIA H100 4.Native hardware acceleration for sparsity (zeros in data). AI models are designed to be over-parameterized and created with more parameters than required to produce the result. The goal of training a neural network is to find the important subset of the parameters to produce the best result. Therefore, AI models can contain a large number of zero-valued parameters, which are less important, and do not impact the result. These zero-valued parameters are considered “sparse” and can mathematically be skipped when computing the model, since multiply-by-zero is always zero. The WSE-3 architecture has the memory bandwidth and built-in fine-grained dataflow hardware mechanisms to skip multiplies by zero, only performing work for the non-zero important parameters, and thereby accelerating performance. We believe the WSE-3 is the only commercially available hardware that can accelerate every sparsity pattern because our hardware supports fully unstructured sparsity. This sparsity acceleration uses less power by skipping unnecessary compute and can provide a significant additional performance advantage both for training and inference. !business14a.jpg

*business14a.jpg*

Sparsity Within Neural Networks. Dense networks can be made sparse by removing the less important parameters (connections between nodes) and retaining only the important information. We believe Cerebras has the only hardware architecture on the market capable of accelerating all forms of sparsity, including unstructured sparsity. 122


CS-3 System Powered by the WSE-3 – Innovative Power and Cooling for Our Wafer-Scale Chip The CS-3 system houses the WSE-3 and is engineered to tackle the unique thermal and power challenges of wafer-scale integration. To support the wafer-scale chip and make it easy to deploy in modern data centers, we have developed and productized key system and packaging technologies: •Thermal expansion-tolerant packaging. The physical connection between the chip and the surrounding infrastructure must be tolerant to the physical stresses caused by thermal expansion over a wide temperature range. At wafer-scale, the thermal expansion stress is significantly higher than traditional chip packages because of the size of the chip. We have developed a unique packaging technique using flexible connections to compensate for the high degree of thermal stresses, designed to enable reliable and high performance in all workload scenarios. •Wafer-scale power and cooling. Since the wafer-scale chip is significantly larger in size than traditional chips, it requires innovative methods of power and cooling that cannot be satisfied by typical methods designed for smaller chips. To address this challenge, we have developed a novel, perpendicular power and water-cooling delivery system designed to provide uniform power and cooling to the entire surface of the wafer, maximizing power efficiency and thermal management. Additionally, the CS-3 provides all the physical infrastructure, interfaces, and management to easily integrate into a standard data center. The CS-3 is a 16RU chassis with high-availability, redundant, and hot-swappable power supplies and fans. The system uses high-speed, standards-based 100G Ethernet connections to communicate to the cluster and rest of the data center. We have also developed an advanced system manager that is constantly monitoring and adjusting the unique wafer-scale operating conditions and providing system administrator interfaces. !business15a.jpg

*business15a.jpg*

Cerebras CS-3 System. The left diagram shows the wafer packaging, called the “engine block” providing direct power and cooling to the wafer. The right diagram shows the entire CS-3 system, which houses a single WSE-3 and engine block, and provides the physical infrastructure to integrate into a standard data center. The Cerebras AI Supercomputer – Purpose Built for Scaling Training and Inference Performance While a single CS-3 already has the equivalent performance of a GPU cluster, there is often a need to accelerate even further by using multiple CS-3 systems. Our specialized wafer-scale AI Supercomputer architecture is designed to scale near-linearly on up to 2,048 CS-3 systems, and to bring exceptional acceleration for training and inference 123


on even the largest AI models, without distribution complexity. Through co-design with the WSE-3 chip and CS-3 system, Cerebras has developed and productized key cluster-scaling technologies: •Training parameter storage. The cluster uses a device called MemoryX to centrally store model parameters during training, enabling seamless model scaling. Parameters are streamed to the CS-3 system for training computations, allowing even a single CS-3 to train the largest models, limited only by MemoryX storage capacity. This separation of storage from compute allows for scaling the model size independent of compute capacity. In the CS-3 wafer-scale cluster, SKUs support models ranging from 30 billion to 24 trillion parameters. •Data-parallel training interconnect. The cluster features a specialized interconnect, SwarmX, enabling data-parallel scaling for training large models. SwarmX includes active devices for weight broadcast and gradient reduction, simplifying operations compared to traditional GPU clusters. In the CS-3 wafer-scale cluster, SwarmX uses 400G and 800G links, designed to support up to 2,048 CS-3 systems in a single cluster. •Inference runtime processing. The AI Supercomputer has proprietary software that handles inference runtime operations. Model parameters are stored in on-chip memory for ultra-low latency during inference. The inference runtime subsystem provides low latency data input to each CS-3, coordinates their operations, and interfaces with external inference systems. It also manages numerous simultaneous inference streams while maintaining ultra-low latency in the CS-3 wafer-scale cluster. !business16a.jpg

*business16a.jpg*

The Cerebras Wafer-Scale Cluster. The single cluster architecture has resources to support training (shown on the left) and inference (shown on the right). These resources are assigned to CS-3 systems (shown in the middle) such that any CS-3 system can be used either for training or inference. CSoft Compiler Co-designed With the Hardware to Deliver Easy-To-Use Performance The Cerebras Software platform (CSoft) is foundational to the usability and user productivity advantages of the Cerebras solution. CSoft allows ML users to program for Cerebras Supercomputer clusters as simply as they would for a single computer, without any distributed computing complexity. This enables rapid iteration and allows users 124


to focus purely on ML experimentation, model quality, and business value creation. To achieve these goals, we developed and productized key software and AI technologies: •PyTorch tracing and integration. PyTorch is the industry-standard programming language of ML practitioners because it is easy to use and abstracts away the hardware. PyTorch allows users to focus on the ML algorithm, programming in the high-level Python language while avoiding low-level languages such as CUDA. CSoft is tightly integrated with PyTorch so that programs written in PyTorch can be extracted and then mapped to the Cerebras architecture efficiently. The PyTorch integration uses an advanced tracing technique that can analyze the AI model prior to executing it, to identify all the model components to be globally optimized. This tracing technique is designed to be leverageable for other ML frameworks that gain future popularity as well, allowing us to adapt as the needs of the ML community change. •Performance optimizing graph compiler. Once the AI model has been extracted from PyTorch, the CSoft Graph Compiler compiles the model for the Cerebras hardware and automatically generates the machine code that runs on each core of each WSE-3. The compile process uses a series of proprietary steps to transform the high-level program to efficient, low-level machine code. To find an optimal mapping, we developed a series of advanced optimization algorithms using constraint solving and simulated annealing to produce high-performance out-of-the-box on the CS-3 hardware. •AI model and workflow libraries. On top of the PyTorch and compiler foundation, the remaining critical enabler to state-of-the-art AI is the ML algorithmic techniques needed to create and run the model for training or inference. We have developed and validated state-of-the-art ML techniques that have led to many research publications and releases of best-in-class models, including GPT model scaling laws, long context architectures, multi-modal image and text models, and hyper-parameter tuning recipes. The list of techniques continues to grow as we further our applied ML research and learnings from customer engagements. We have distilled these techniques into a library of AI models, layer abstractions, and workflows that enable users to directly achieve state-of-the-art ML results, without having to build the model or learn the techniques from scratch. This library is called the Cerebras Model Zoo, an open-source repository of reference model implementations using the latest ML techniques, data preprocessing tools, and training recipes. Model Zoo includes a wide range of GenAI model examples supporting a diverse set of downstream applications and provides users with an easy foundation to build upon that already leverages state-of-the-art advancements and best practices. !business17a.jpg

*business17a.jpg*

The CSoft Software Stack. User programs are extracted from industry standard PyTorch, optimized and lowered to execute on the CS-3 systems with high performance without further user intervention. 125


Extraordinary Performance, Efficiency, and Ease of Use – Fundamental Technology Innovation Produces Industry-Leading User Value The above Cerebras technology advancements combine to provide direct customer benefit across key areas, solving fundamental challenges faced by GPU clusters: 1.High performance training: over 10 times faster training time-to-solution than the leading 8-way GPU systems of the same generation (as reported by many of our customers), designed to scale to 256 exaFLOPS. 2.High performance inference: over 10 times faster output generation speeds than GPU-based solutions from top CSPs. 3.High power efficiency: three times higher performance per watt than the leading 8-way GPU system. 4.Easy to program, with low developer switching cost. Sales and Marketing Our sales and marketing strategy centers on deep market understanding and customer-centric product development. We leverage our extensive market knowledge, proven track record in delivering large-scale compute solutions, and close customer collaborations to optimize our product roadmap. This is designed to ensure our solutions consistently deliver significant value to our customers. Alongside a focused sales force, we maintain a dedicated Field ML team. This team provides customers with access to leading AI expertise, ensuring they are positioned to leverage our technology effectively. Field ML teams are supported by Applied ML teams, product applications engineers, marketing, and business development/strategy teams, fostering a comprehensive customer support structure. We focus our sales and marketing efforts on industry leaders, specifically large enterprises domestically and abroad with rich data assets. Our customers are seeking to leverage their rich proprietary data and combine it with Cerebras’ industry leading compute and AI expertise to build a durable competitive advantage. Among our customers, word of success travels quickly, and as a result, it is very important to our future that we maintain strong and collaborative relationships and that we invest behind the largest and most successful of our customers. We utilize master purchase agreements, purchase orders, and statements of work, to define work scope, price, quantities, delivery terms, warranties, and software subscriptions. We predominantly sell our solutions directly to customers via on-premise hardware or via the cloud, based on a consumption-based model. Research and Development We are committed to relentless innovation in both hardware and software to address the rapidly-evolving computational needs of GenAI. In the hardware domain, our research and development efforts focus on chip and system design. We spearheaded the development of wafer-scale integration, resulting in three generations of industry-leading processors at the 16nm, 7nm, and 5nm process nodes. While the idea that bigger silicon would lead to computational efficiencies was novel in 2016, the rest of the industry has now moved in this direction. Cerebras’ innovations in wafer-scale technology helped solve a technical problem that had been open in the processor industry for 75 years. Core members of our founding and early engineering team are world-leaders in chip design and took big risks to help drive the computational industry forward. Our rigorous design methodology utilizes cutting-edge simulation tools and partnerships with Electronic Design Automation providers to create high-quality processors and accelerated development cycles. 126


In the software domain, we develop specialized compilers that translate AI models for optimal performance on our hardware, maximizing performance and efficiency. Our researchers are leaders in sparsity techniques, a critical approach for efficient AI workloads, as evidenced by our publications in this field. These techniques are represented in new and advanced models that we make available to our customers. We dedicate significant resources to ongoing research and development. We invest heavily in attracting and retaining a global team of highly skilled engineers across dedicated facilities in the United States, Canada, and India. This unwavering commitment to innovation fuels our growth and positions us as a leader in the GenAI landscape. Manufacturing and Suppliers We operate a fabless manufacturing model, strategically partnering with industry leaders for the production of our AI compute systems which include ICs, boards, and systems. Our core manufacturing partners include: TSMC, a leading semiconductor foundry, fabricates our cutting-edge WSEs. Advanced Semiconductor Engineering (“ASE”) handles specialized processes, including the deposition of redistribution layers, and we manage final wafer packaging, assembly, and testing in our Sunnyvale, California facility. We also use a small number of third parties to manufacture subassemblies and critical components such as printed circuit boards, I/O subsystems, cooling assemblies and power delivery modules. The manufacturing process is subject to extensive testing and verification. Our supply chain is designed for flexibility and for quality, as we plan to ramp up production to meet the growing global demand for our AI compute systems. Simultaneously, we are committed to rigorous quality control throughout the manufacturing process to confirm reliability in even the most demanding environments at our customer facilities. Our contract manufacturing partners perform system assembly, extensive testing, and verification protocols are in place at every stage including post assembly. Final system-level burn-in and test is conducted by Cerebras. Our quality processes include high production test coverage, full product traceability, and extensive post assembly burn-in. We employ a dedicated quality team that continuously monitors feedback during manufacturing and after deployment. This data-driven approach allows us to improve our product quality and reliability, and help us meet the stringent demands of our customers worldwide. Intellectual Property Protecting our intellectual property and proprietary technology, including our AI products and solutions, is an important aspect of our business. We rely on a combination of intellectual property rights, including patent, trademark, trade secret, and other related laws in the United States and internationally as well as confidentiality procedures and contractual provisions to protect, maintain, and enforce our proprietary technology, intellectual property rights, and brand. Our intellectual property portfolio includes patents, trademarks, proprietary software, and trade secrets. As of June 30, 2024, we owned 85 issued patents and 15 pending patent applications globally. Of these, 44 are issued U.S. patents and nine are pending U.S. patent applications. Our issued patents and pending patent applications generally relate to the design and fabrication of wafer-scale processors, the assembly, packaging, and cooling of wafer-scale processors and hardware, and software architectures for accelerated deep learning. The expiration dates of the U.S.-issued patents are between 2038 and 2042, not taking into account any applicable patent term extensions. We routinely review our development efforts to assess the existence and patentability of new inventions. We have a policy of requiring employees and consultants to execute confidentiality agreements upon the commencement of an employment or consulting relationship with us. Our employee and independent contractor agreements also require relevant employees and independent contractors to assign to us all rights to any inventions made or conceived during their employment or engagement with us. In addition, we typically require individuals and 127


entities with which we discuss potential business relationships to sign non-disclosure agreements that contain customary confidentiality provisions. Competition We offer a purpose-built AI compute platform. Our CS product family primarily competes against solutions from NVIDIA Corporation, Advanced Micro Devices, Inc., Intel Corporation, Microsoft Corp., and Alphabet Inc., among others, as well as internally developed custom application-specific integrated circuits and a variety of private companies, some of which are focused on inference-only offerings. We believe that our ability to remain competitive will depend on how well we are able to anticipate the features and functions that customers will require and whether we are able to deliver consistent volumes of our products at acceptable levels of quality and at competitive prices. We expect competition to increase from both existing competitors and new market entrants with products that may be lower priced than ours or may provide better performance or additional features not provided by our products. In addition, it is possible that new competitors or alliances among competitors could emerge and acquire significant market share. Some of our competitors have greater marketing, financial, distribution, and manufacturing resources than we do and may be more able to adapt to customers or technological changes. We expect an increasingly competitive environment in the future. Human Capital As of September 26, 2024, we had 401 employees, including 252 in the United States, and we have employees located internationally, including in Canada, and in India. We maintain a full-time workforce and supplement our workforce with contractors and consultants. To our knowledge, none of our employees are represented by a labor union or party to a collective bargaining agreement. We consider our relationships with our employees to be good. Our human capital resources objectives include, as applicable, identifying, recruiting, retaining, incentivizing, and integrating our existing and new employees. The principal purposes of our equity incentive plans are to attract, retain, and reward personnel through the granting of stock-based compensation awards in order to increase stockholder value and the success of our company by motivating such individuals to perform to the best of their abilities and achieve our objectives. Facilities Our corporate headquarters is located in Sunnyvale, California, where we lease approximately 68,000 square feet for office space, research and development, and testing, pursuant to a lease agreement that expires in November 2027, subject to the terms thereof. We lease additional facilities in San Diego, Canada, and India for research and development. We also enter into agreements for offsite colocation facilities to house and operate our AI supercomputers. We enter into these agreements for our own corporate purposes as well as on behalf of our customers. Currently, our data center facilities are in California, and we manage another data center facility in Texas on behalf of a customer. We believe that our facilities are suitable to meet our current needs. We intend to expand our facilities or add new facilities as we grow, and we believe that suitable additional or alternative spaces will be available on commercially reasonable terms, if required. Government Regulations We are subject to many U.S. federal and state laws, rules, and regulations, as well as laws, rules, and regulations imposed by various non-U.S. governmental authorities, including those related to intellectual property, tax, import and export requirements, anti-corruption, economic and trade sanctions, national security and foreign investment, foreign exchange controls and cash repatriation restrictions, data privacy and security requirements, competition, advertising, employment, product regulations, environment, health, and safety requirements, and consumer laws. 128


These laws and regulations are complex, are constantly evolving, and may be interpreted, applied, created, or amended, in a manner that could harm our business. The import and export of our products and technology are subject to laws and regulations, including international treaties, U.S. and various non-U.S. export controls and sanctions laws, customs regulations, and other trade rules. The scope, nature, and severity of such controls varies widely across different countries and may change frequently over time. Such laws, rules, and regulations may delay the introduction of some of our products or impact our competitiveness through restricting our ability to do business in certain countries or territories or with certain parties (including certain governments) or certain jurisdictions. U.S. export restrictions also require us to obtain licenses from the U.S. Department of Commerce to allow the export or transfer of our products (including our software and technology), and there can be no assurance that export permissions will be granted. See the section titled “Risk Factors” for additional information regarding risks we face related to government regulation. Legal Proceedings From time to time, we may be subject to legal proceedings, claims, and investigations in the ordinary course of business. We are not presently a party to any litigation to which the outcome, we believe, if determined adversely to us, would individually or taken together have a material adverse effect on us. We cannot predict the results of any such proceedings, claims, or investigations, and despite the potential outcomes, the existence thereof may have a material adverse impact on us due to diversion of management time and attention as well as the financial costs related to resolving such matters. 129