Internal tensions ran high at Amazon Web Services (AWS) in early 2025 as its newly launched AI service, Bedrock, suffered from frequent technical glitches and performance bottlenecks that frustrated customers.
According to sources familiar with the matter, users repeatedly encountered error messages and often waited weeks to receive the computing power they had requested. In response to these mounting issues, Anthony Ligori, a seasoned engineering lead, assembled a team of five other senior engineers to devise a comprehensive overhaul of Bedrock — an initiative code-named Project Mantle.
The project aimed to resolve rate-limiting and error issues while expanding the service's capacity to accommodate a growing user base. Ligori and his colleagues leveraged AWS's proprietary AI coding tool, Kiro, to build Mantle, which officially launched in December 2025.
Shortly after its release, AWS added OpenAI models to Bedrock's portfolio, which already included Anthropic's Claude and other leading models. The platform's popularity surged almost immediately. Randy Hunt, Chief Technology Officer at consulting firm Caylent, noted that some clients have already shifted their cloud budgets from Microsoft Azure to AWS.
Although AWS holds a larger share of the cloud market and generates higher revenue than Azure, Microsoft had previously maintained a competitive edge by being the sole cloud provider hosting OpenAI models. Hunt observed that clients now benefit from reliable, consistent performance on Bedrock, whereas Azure's performance remains unpredictable. He added that his financial-sector clients have quadrupled their AWS spending since May, and in his assessment, AWS's AI service outperforms rivals like Microsoft.
Amazon disclosed in July that Bedrock added more new customers in the first half of 2026 than in its first two years combined, with second-quarter customer spending exceeding all previous quarters total. While Amazon does not separately disclose Bedrock's revenue, its financial reports show AWS's second-quarter growth accelerating by 9 percentage points to 37%, compared to Azure's 42% growth, which rose 2 percentage points quarter-over-quarter. Microsoft stated that as AI adoption expands, workloads are shifting from model development to large-scale inference and agent-based applications, and the company continues to evolve its AI stack to meet changing customer demands and growing compute needs.
Bedrock's early struggles highlight how even the world's largest cloud provider can face significant challenges from the immense compute demands of AI. An AWS spokesperson said the company rewrote Bedrock's inference engine from scratch to keep pace with surging demand, balancing performance and security, and the approach has been widely embraced by customers. The spokesperson described Bedrock as AWS's fastest-growing multi-billion-dollar service, with hundreds of thousands of customers.
Technical hurdles and early missteps
Following ChatGPT's release in late 2022, which ignited the AI race, AWS rolled out Bedrock in fall 2023. Despite already offering AI tools like SageMaker, Amazon initially struggled to gain momentum in the AI wave, creating an opening for Microsoft. One project participant noted that AWS engineers worked 60-hour weeks to build the service that allows customers to connect cloud applications to large language models. The service was partly driven by customer demand for running leading AI models that could not operate on existing systems. Unlike standard software, large models utilize proprietary weights tightly guarded by AI companies, with model vendors prohibiting direct customer access to these weights.
However, the hastily developed Bedrock left a sense of patchwork quality, according to two later project participants. An insider and another team member said the original architecture failed to fully leverage AI chip processing power, causing frequent customer errors and limiting Bedrock's capacity to scale broadly. An AWS spokesperson acknowledged that the company continually iterates on its services to achieve perfection, but occasional unforeseen issues arise at AWS's scale, which the team identifies, communicates, and resolves as quickly as possible.
Simultaneously, customer compute quota problems intensified. Each user received a baseline allocation of computing resources, and once usage exceeded that limit, they were throttled and had to wait for AWS to release more capacity. A former employee reported wait times of up to three weeks for some customers. A product team member revealed that throughout 2025, customers repeatedly complained about restrictions on accessing Anthropic models. AWS management even worried that AI coding startup Lovable might leave the platform due to difficulties invoking Claude through Bedrock. A person who viewed an internal document from early 2025 stated that AWS's sales team circulated a report highlighting how Bedrock's technical flaws were hampering customer acquisition. Around the same time, The Information reported that executives described Bedrock's compute crisis as a disaster.
Hunt noted that in 2024, customers testing Bedrock found throughput for Anthropic models was up to 63% slower than connecting directly to Anthropic. Businesses connect directly to model providers through APIs, though models remain hosted on cloud servers; Anthropic also offers direct API access through Google Cloud alongside AWS. Bedrock's problems even drew criticism from leaders in other Amazon business units. Two people who heard the conversations said Dave Treadwell, then senior vice president of Amazon's e-commerce division, told AWS management in early 2025 that Bedrock's persistent errors and compute shortages made him want to switch directly to OpenAI's native models, which are primarily hosted on Microsoft Azure. A person close to Amazon disputed Treadwell's characterization; Treadwell has since moved to AWS.
The Mantle solution takes shape
Ligori then proposed building a new version of Bedrock. At AWS's annual re:Invent customer conference in December 2025, Dave Brown, then head of the cloud server business, unveiled Mantle — a new technology layer powering Bedrock designed to enable continuous scaling to meet anticipated future demand. During his keynote address, Brown explained that as models evolve and usage rises, the team recognized the need for a new architecture for inference. (Brown has since left AWS to join Meta.)
Brown also pointed to the original Bedrock's architectural flaws: it was designed following the blueprint of traditional web services. However, running large models that drive applications — a process known as inference — is far more complex than conventional cloud workloads. Traditional cloud services mainly involve customer storage, retrieval, and data analysis, whereas inference requires dynamically provisioning sufficient servers to handle highly variable loads. This discrepancy became increasingly problematic in 2025 as customers began using AI for more sophisticated tasks. Brown stated at re:Invent that inference operates fundamentally differently from the traditional computing models AWS has optimized for the past two decades.
Mantle specifically addresses the spiky compute demands of complex AI tasks, where processing needs surge dramatically and then quickly recede. Early AI tasks were relatively simple and could be supported by the old architecture, but new long-running AI agent tasks require a distinct compute allocation mechanism, as an AWS spokesperson explained. To handle compute spikes at scale, Mantle introduced a feature allowing customers to set priorities for different tasks within Bedrock, enabling AWS to direct more computing resources to high-priority work. Mantle also includes load isolation, which the spokesperson compared to a concert venue having multiple independent entry points — a single customer's heavy load only impacts their own allocation without crowding out other users' resources.
AWS also enhanced Mantle's compatibility with its existing Journal service, which saves the progress of AI agent tasks so long-running operations can resume without restarting from scratch, accommodating the increasing prevalence of extended AI workflows. Joe Maglalamov, a former AWS vice president who worked on Mantle, noted on an AWS podcast in January that through Mantle, the team realized inference is not fundamentally a web service but rather more akin to a scheduling system. (Maglalamov joined Meta last month following Brown.)
Native API compatibility and market response
Mantle achieves direct compatibility with OpenAI's and Anthropic's native APIs. The original Bedrock required customers to write custom code to interface with these AI labs' model APIs. Mantle was fully deployed in early 2026. A former employee said customer wait times for capacity expansion shrank from three weeks in early 2025 to just two days by year-end. In his 2026 letter to shareholders, Amazon CEO Andy Jassy called Mantle the technological foundation of the highly successful Bedrock service.
Amazon stated that starting in late 2025, Bedrock largely ran on its self-developed AI training chip, Trainium, providing AWS with an additional compute source amid ongoing constraints in Nvidia chip supply. Of course, customers still occasionally experience latency with the service. One business manager using Mantle to invoke Claude on Bedrock reported intermittent service disruptions over recent months, with the longest outage lasting about 11 hours. Bojan Jakimovski, machine learning lead at AI consulting firm Loka, said six clients have already moved to running OpenAI models on Bedrock instead of connecting directly to OpenAI or using Azure. Jakimovski noted the migration brings multiple advantages, including faster inference, higher overall throughput, and unchanged costs with all other parameters identical.
Comments