Stimuler

Case Study

Stimuler Powers Real-Time English Coaching for 15 Million Users on Anyscale

Stimuler runs a composite AI system that combines four specialized models for English speech coaching. Ray and Anyscale let a four-person team scale each one independently across 900,000 daily sessions.

stimuler-case=study-hero-image1 v2

15M+

users served across Indonesia and Spanish-speaking Latin America

900K+

speeches processed daily across a distributed multi-model pipeline

4

person ML team operating a production-scale compound AI system on AWS

Stimuler is an AI-native English speaking platform purpose-built for non-native speakers in high-growth markets. Founded in 2023 and now serving more than 15 million users across Indonesia and Spanish-speaking Latin America, Stimuler's core product is a voice-first coaching experience in which an AI avatar acts as a personalized English teacher: listening to each utterance, assessing proficiency in real time, deciding the ideal pedagogical response, and speaking back to the learner in a natural, adaptive way. The company traces its origins to a viral moment in Indonesia, where a screenshot of a user's practice score spread across Twitter and brought 10,000 new signups in a single day, providing the escape velocity the founders needed while still in college.

What makes Stimuler technically demanding is the nature of its user base. Speakers from Indonesia and Latin America arrive with strong first-language influence on their English, mixing Bahasa or Spanish with English mid-sentence and producing phoneme patterns that differ sharply from what standard speech-to-text systems are trained on. The goal is not to infer intent as most automatic speech recognition (ASR) systems do, but to capture exactly what was said so the platform can analyze mispronunciation at the phoneme level, track proficiency changes after every utterance, and provide correction grounded in how each learner's specific native language shapes their spoken English. Delivering that at the scale of 900,000 speeches per day, across four independent model stacks, required infrastructure that could scale each component on its own terms without bottlenecking the others. Stimuler chose Ray and Anyscale to make that possible.

LinkChallenges

Building a real-time speech coaching platform for non-native English speakers at consumer scale means solving a composite AI problem: multiple specialized models serving different functions, each with its own resource profile and latency requirement, all needing to run concurrently and scale independently to serve a single request without a large engineering organization behind them. 

Three challenges stood in the way:

  • Iterating on four independent model stacks with a four-person ML team required a platform where experimentation and production felt like the same workflow. Stimuler's platform spans a speech understanding stack, an assessment engine, a proficiency tracking system, and an interaction stack for speech generation and decision-making. Each of those stacks contains multiple models, and in some cases multiple adapters mounted on a shared base, all of which the ML team is continuously improving. Evaluating a new demographic-specific speech model, benchmarking an updated adapter configuration, and promoting a validated candidate to production are all part of a regular development cycle. Doing that across four stacks with four engineers meant any friction between the experimentation environment and the production environment would directly slow the pace of model improvement, which is the core driver of product quality for Stimuler.

  • A compound AI system with four independent stacks needed to scale each component separately without bottlenecking the others. Each part of Stimuler's pipeline receives requests at a different rate and carries a different resource profile. Proficiency tracking fires after every utterance; speech generation fires only when the AI teacher responds. Some components require GPUs for large analytical models handling transcription and multi-task inference; others require fast CPU-bound models delivering instantaneous highlighting and voice-activity feedback over WebSocket connections. A platform that forced all stacks to share a single scaling policy would either overprovision resources across the board or create the kind of cascading latency that breaks real-time user experience. The infrastructure had to treat each model stack as an independently scalable unit while still coordinating them within a single coherent pipeline.

  • Running production infrastructure at the scale of millions of users with no dedicated platform engineering function carried a hidden engineering cost that had to be designed around. Stimuler's ML team of four is responsible for building, evaluating, and deploying custom speech models across multiple demographic distributions simultaneously. Every hour spent provisioning clusters, debugging autoscaling behavior, or recovering from infrastructure failures is an hour not spent on the speech modeling and pedagogical improvements that drive user outcomes. At 900,000 speeches per day, the operational surface of the production system is significant, and sustaining it without growing the team required a managed platform that absorbed the infrastructure burden automatically rather than passing it back to the people building the models.

"Our platform has to support multiple downstream AI tasks simultaneously, from transcription to assessment and interactive learning experiences. Different user populations require different speech models, and having a single platform that could serve all of those workloads without separate infrastructure for each was essential."
Ankit Kumar Pandey's profile

Ankit Kumar Pandey | Co-founder, Stimuler

Stimuler logo

LinkThe Solution

Both co-founders started Stimuler with experience in distributed systems, and when the architecture required independent scaling across four model stacks with heterogeneous compute, Ray emerged as the natural fit. The team adopted Anyscale to handle the operational layer on top of Ray, deploying and autoscaling clusters on AWS EC2 without dedicating engineering time to maintaining that infrastructure themselves. Stimuler was part of the AWS Generative AI Accelerator cohort of 2025, one of 40 startups selected for the program, and Anyscale played a direct role in helping the team stand up and scale their Ray deployments through that engagement.

With Anyscale, Stimuler is able to:

  • Run the full model development lifecycle, from experimentation to production, through a single platform. The ML team uses Anyscale Workspaces for experimentation and training and Anyscale Jobs for benchmarking, giving every stage of model development a consistent interface so the team can move between exploration and production without rebuilding workflows or switching environments.

  • Scale four independent model stacks without coupling their resource profiles. Each component of the speech coaching pipeline can grow or contract based on its own demand curve, with CPU-bound models handling real-time feedback workloads and GPU-bound models handling transcription, proficiency updates, and multi-task inference, all coordinated through Ray's deployment and service abstractions without separate infrastructure for each.

  • Operate production-scale distributed infrastructure without the engineering overhead of managing it. Anyscale handles cluster provisioning, autoscaling, and operational maintenance on AWS EC2, freeing the ML team to focus on the speech modeling and pedagogy problems that differentiate the product rather than on the infrastructure layer beneath them.

"Anyscale has been instrumental in deploying, maintaining, and autoscaling Ray deployments for us. Our whole way of building is intertwined with Anyscale, and we feel comfortable depending on the platform."
Ankit Kumar Pandey's profile

Ankit Kumar Pandey | Co-founder, Stimuler

Stimuler logo

LinkA unified platform for experimentation and production

With four ML engineers responsible for building, evaluating, and deploying custom speech models across multiple demographic distributions simultaneously, the cost of context-switching between experimentation and production infrastructure was a real constraint on development velocity. Stimuler's model development cycle is always evaluating new demographic-specific speech models, iterating on adapter configurations, and benchmarking candidates against prior versions across multiple language profiles at once. Any friction at the boundary between experimentation and production would compound across all of that activity, eating directly into the time available for the model improvements that drive product quality.

Anyscale Workspaces serve as the primary environment for model experimentation and training at Stimuler, with Anyscale Jobs handling the benchmarking layer, giving every stage of the model lifecycle a consistent interface grounded in the same Ray abstractions that govern the production deployment. A model that clears benchmarks in a Job and shows the right behavior in a Workspace can move toward production without a re-architecture step at the boundary, removing the context-switching overhead that would otherwise force the team to maintain separate toolchains for exploration and deployment.

That continuity across the full lifecycle means the three-person ML team can operate with a development cadence that would typically require a much larger organization to sustain, resulting in faster iteration across four stacks and multiple demographic variants than a fragmented toolchain would allow.

"Jobs and Workspaces became an important part of how we benchmark, experiment, and deploy models with a very lean ML organization. Anyscale lets us focus on building better speech intelligence instead of managing infrastructure."
Ankit Kumar Pandey's profile

Ankit Kumar Pandey | Co-founder, Stimuler

Stimuler logo

LinkIndependent scaling across a compound AI system

Stimuler's coaching pipeline is not a single model with a single throughput curve. Four distinct stacks, each drawing on multiple models with different request volumes and resource profiles, needed to scale independently without any one component creating a bottleneck for the others. A user session touches all four in sequence and in parallel: speech understanding captures and analyzes the utterance at the phoneme level, the assessment engine evaluates content and grammar, proficiency tracking updates the learner's evolving model, and the interaction stack decides the next action and generates the teacher's spoken response. Each fires at a different rate and on a different resource budget, and none can afford to wait on the others without degrading the real-time experience the product depends on.

Ray's deployment, service, and application abstractions give the ML team the ability to define each component as an independently scalable unit, meaning a proficiency tracking model that fires on every utterance scales on a different curve than the speech generation model that fires when the teacher responds, without either influencing the resource allocation of the other. Where managing this kind of heterogeneous, multi-service architecture would otherwise mean maintaining containers and Kubernetes configurations for each compute class separately, Ray's Python-native abstractions let the team define, deploy, and scale each stack without leaving the ML development environment they already work in. Voice-activity detection and instantaneous reading-highlight models run on CPUs over WebSocket connections, while the larger analytical models for transcription, mispronunciation scoring, and proficiency updates run on GPUs, with Anyscale coordinating both within the same deployment.

"Each part of the system has multiple models and adapters, and each needs its own scaling based on its own request pattern. Ray and Anyscale are the perfect fit because of distributed compute and the concepts of deployments, services, and applications, where you can abstract and scale different things independently."
Ankit Kumar Pandey's profile

Ankit Kumar Pandey | Co-founder, Stimuler

Stimuler logo

LinkProduction infrastructure at scale without a platform team

For a four-person ML team with no dedicated platform engineering function, the operational cost of running a production AI system is massive. Every hour spent provisioning clusters, debugging autoscaling behavior, or recovering from infrastructure failures is an hour taken directly from the speech modeling and pedagogical work that determines how well the product serves its users. Sustaining a system of Stimuler's complexity without growing the team required a platform that could handle interruptions, scaling events, and cluster maintenance without requiring the engineers to step in.

Stimuler runs its full production stack within their own AWS account through Anyscale, relying on it to handle cluster provisioning, autoscaling, and operational maintenance while the ML team stays focused on model development. Anyscale's managed Ray clusters absorb all of that automatically since the company's early days, freeing a three-person team to process more than 900,000 speeches daily across heterogeneous CPU and GPU workloads without the platform headcount a system at that scale would typically demand.

"We have been working with Anyscale for more than a year now and our whole way of building is kind of intertwined with Anyscale. It has helped us deploy the models within our own AWS account and scale on AWS EC2, which helps us utilize and work with that very efficiently."
Ankit Kumar Pandey's profile

Ankit Kumar Pandey | Co-founder, Stimuler

Stimuler logo

LinkWhat's Next

Stimuler is expanding its model stack to include text-to-speech and multimodel output models built on open-source Omni foundation models, with multiple task-specific adapters serving different demographics and use cases. As the platform grows at roughly five to six percent monthly and the model architecture becomes more consolidated, the distributed compute requirements will continue to scale with it, and Anyscale is expected to grow alongside them.

"Ray and Anyscale gave us the flexibility to scale different parts of our speech understanding platform independently without adding operational complexity."
Ankit Kumar Pandey's profile

Ankit Kumar Pandey | Co-founder, Stimuler

Stimuler logo

“Ray and Anyscale fit perfectly into the architecture we needed to power speech understanding at global scale. With a four-person ML team, we are able to operate AI systems serving millions of learners worldwide."

Ankit Kumar Pandey

Co-founder, Stimuler

ankit-kumar-pandey-co-founder-stimuler
Want to give it a try?