Replit turned a browser-based code editor into an AI-powered app factory, and it did it without training a single large language model of its own. That surprises a lot of developers. When people ask what LLM does Replit use, they often expect to hear about some secret in-house model humming away in a data center. The reality is more interesting: Replit built its reputation on smart orchestration of the best third-party models available, mixing Anthropic’s Claude family, OpenAI’s GPT series, and Google’s Gemini models depending on the job at hand.
That choice shapes everything about how Replit Agent behaves, how much your builds cost, how fast your code appears, and how reliably your app actually runs. In this guide, you’ll learn exactly which models Replit leans on today, how the company routes requests between them, why Claude became the backbone of Replit Agent, what happened to Replit’s own homegrown code model, and how Replit stacks up against competitors like Cursor, Lovable, and GitHub Copilot. You’ll also get practical tips for getting better results, avoiding common mistakes, and understanding where Replit’s AI stack is heading next.
The Models Powering Replit’s AI Tools
Replit does not train its own frontier large language model anymore; instead, Replit Agent and Replit Assistant run primarily on Anthropic’s Claude models, with Claude Sonnet serving as the main engine, supported by OpenAI’s GPT models and Google’s Gemini models for specific tasks. Replit’s leadership has been open about this. CEO Amjad Masad has publicly credited Anthropic’s Claude Sonnet releases as the turning point that made Replit Agent viable as a product, and Replit has appeared repeatedly in Anthropic’s customer case studies as one of the largest consumers of Claude API tokens for coding workloads.
Think of it like a restaurant kitchen. Replit is the chef and the kitchen design. The LLMs are the ingredients suppliers. Replit’s real product isn’t the model itself, it’s everything wrapped around the model: the file system access, the package installer, the database provisioning, the deployment pipeline, the error checker that reads your app’s logs and fixes bugs on its own. The model provides raw reasoning and code generation. Replit provides the arms and legs that turn that reasoning into a running application.
This matters for you as a user because it means Replit’s capabilities jump whenever Anthropic, OpenAI, or Google ship a better model. When Claude Sonnet 3.5 landed, Replit Agent got noticeably smarter. When Claude Sonnet 4 and 4.5 arrived, Agent 3 became capable of running for hours without a human babysitting it. Replit typically adopts new flagship models within days or weeks of release, which is far faster than any company building models from scratch could iterate.
Here’s a quick snapshot of the model landscape inside Replit:
| Model Family | Provider | Typical Role in Replit |
|---|---|---|
| Claude Sonnet (3.5, 3.7, 4, 4.5) | Anthropic | Primary engine for Replit Agent code generation, planning, and multi-step builds |
| Claude Haiku | Anthropic | Fast, cheap tasks like small edits, classification, and routing decisions |
| Claude Opus | Anthropic | High-difficulty reasoning and complex architectural decisions when available |
| GPT-4o / GPT-4.1 / GPT-5 class | OpenAI | Alternative reasoning, certain Assistant features, fallback capacity |
| Gemini (Flash and Pro tiers) | Long-context tasks, cost-efficient background work, image and design understanding | |
| Embedding models | Mixed vendors | Code search, retrieval, and context selection across your project files |
Why Claude Became Replit’s Model of Choice
Plenty of strong coding models exist, so why did Replit lean so heavily on Anthropic? The answer comes down to three things: instruction following, tool use, and agentic stamina.
Instruction following matters because Replit Agent doesn’t just write a snippet and stop. It receives a long system prompt full of rules about file structure, framework choices, security practices, and formatting. A model that drifts away from those rules produces broken projects. Claude models have consistently scored well on strict instruction adherence, which means fewer moments where the Agent ignores your stack preference and installs something random.
Tool use is the second pillar. Replit Agent works by calling tools: read this file, write that file, run this shell command, search the web, query the database, restart the server. Claude’s function-calling reliability, especially from Sonnet 3.5 onward, made these chains dependable enough to ship to non-technical users. A model that hallucinates a tool call or mangles JSON arguments breaks the whole loop.
Agentic stamina is the newest factor. Replit Agent 3 markets itself on the ability to work autonomously for extended stretches, testing its own output in a browser, spotting failures, and repairing them. That requires a model that stays coherent across dozens or hundreds of turns without losing the thread. Claude Sonnet 4 and 4.5 raised that ceiling significantly.
What Replit Says Publicly
- Replit has featured in Anthropic case studies describing dramatic growth in Agent usage after adopting newer Claude Sonnet versions.
- Amjad Masad has repeatedly credited Claude Sonnet 3.5 as the model that made the original Replit Agent launch possible in September 2024.
- Replit’s changelog and documentation reference Claude models by name when describing Agent behavior and effort levels.
- Replit also participates in early access programs, which is why new Anthropic models often appear inside Replit within days of announcement.
None of this means Replit is locked in. The company has been deliberate about staying model-agnostic under the hood so it can swap providers if the quality or pricing balance shifts.
How Replit Routes Requests Between Different Models
Replit doesn’t send every keystroke to the most expensive model available. That would burn money and slow everything down. Instead, it uses a routing layer that matches the task to an appropriate model tier. Understanding this helps explain why some Agent actions feel instant while others take minutes.
Here’s roughly how a single Agent request flows through the system:
- You type a prompt describing what you want built or changed.
- A lightweight model or classifier interprets the request and decides whether it needs a quick edit or a full planning cycle.
- For bigger jobs, a planning step runs on a stronger reasoning model, producing a checklist of tasks and files to touch.
- Retrieval kicks in, pulling relevant files, dependency lists, and error logs into context so the model doesn’t guess.
- The main coding model generates edits and issues tool calls to write files, install packages, and run commands.
- A verification pass checks logs, runs the app, and sometimes drives a headless browser to confirm the feature works.
- If something breaks, the loop repeats with the error message included, often on a cheaper model first before escalating.
Consider a practical example. You ask Replit Agent to “add Stripe checkout to my store and email a receipt.” The system doesn’t send that as one giant prompt. It breaks the work into installing the Stripe SDK, creating an API route, adding environment variables through the Secrets manager, wiring the frontend button, connecting an email provider, and testing the flow. Small mechanical steps like renaming a variable might run on a fast, inexpensive model. The architectural decision about where to put the webhook handler runs on a stronger one.
This layered approach is why Replit’s pricing shifted toward effort-based billing. You’re not paying per message anymore; you’re paying for the compute the Agent actually consumed, which depends heavily on which models it called and how many times.
Replit Agent Versus Replit Assistant: Different Tools, Different Model Needs
Replit ships two distinct AI experiences, and people frequently mix them up. Knowing the difference helps you choose the cheaper, faster option when you don’t need the heavy machinery.
Replit Assistant
Assistant handles targeted changes. You highlight a component, ask it to fix a bug, adjust styling, or explain a function, and it makes a focused edit. Because the scope stays small, Assistant can lean on faster and cheaper models with tighter context windows. Responses arrive in seconds, and costs stay low.
Replit Agent
Agent builds and rebuilds entire applications. It creates project scaffolding, provisions databases, sets up authentication, installs packages, writes tests, and deploys. It needs long-context reasoning, reliable tool calling, and the ability to recover from its own mistakes. That means the strongest available models and a lot more tokens.
| Feature | Assistant | Agent |
|---|---|---|
| Best for | Small edits, explanations, quick fixes | Full apps, multi-file features, integrations |
| Model tier | Faster, lower-cost models | Frontier reasoning models |
| Autonomy | Waits for your approval on each change | Runs long autonomous loops |
| Typical speed | Seconds | Minutes to hours |
| Relative cost | Low | Higher, effort-based |
| Can deploy | Limited | Yes, end to end |
A smart workflow uses both. Let Agent build the first version of your app, then switch to Assistant for the twenty small tweaks that follow. You’ll spend far less and get faster iterations.
What Happened to Replit’s Own Code Model
Replit did train its own model once, and the story explains a lot about the current strategy. In 2023, Replit released replit-code-v1-3b, an open-source 2.7 billion parameter code completion model trained on permissively licensed code from the Stack dataset. It supported around 20 programming languages and was designed for fast autocomplete inside the Replit editor. Replit later released a v1.5 version trained on more data.
At the time, this made sense. Frontier models were expensive, slow, and not great at code. A small, specialized model that ran cheaply could handle inline completions well enough. Replit also partnered with Google Cloud for training infrastructure and used it to power features like Ghostwriter, the predecessor to today’s Assistant.
Then the landscape shifted fast. General-purpose frontier models blew past small specialized code models on almost every benchmark, and they added something a 3B model could never do: multi-step reasoning and reliable tool use. Training and maintaining a competitive frontier model costs hundreds of millions of dollars. Replit made the pragmatic call to stop competing on model training and compete instead on the product layer.
The lesson generalizes. Most AI application companies today are not model companies. They are orchestration companies. Their moat comes from the environment they build around the model, the data they collect about what works, and the user experience they design. Replit’s moat is the cloud development environment, the instant deployment, the database provisioning, and the ability to go from idea to live URL without touching a terminal.
That said, Replit still trains smaller specialized models for narrow internal tasks like code repair suggestions, ranking, and classification. Those don’t get headlines, but they quietly improve the product.
Common Misconceptions About Replit’s AI Stack
Confusion around Replit’s models leads people to make bad assumptions about performance, privacy, and cost. Let’s clear up the big ones.
- Misconception: Replit built its own ChatGPT competitor. It didn’t. Replit builds the agent harness, not the frontier model. The reasoning comes from Anthropic, OpenAI, and Google.
- Misconception: You can freely pick any model you want. Unlike some IDEs that offer a model dropdown with a dozen choices, Replit deliberately abstracts model selection away. You control effort level and mode, not the specific model in most cases.
- Misconception: Replit is just a wrapper. The scaffolding is genuinely hard. Sandboxed execution, dependency resolution, secret management, database provisioning, browser-based testing, and one-click deployment represent years of engineering that no model provides on its own.
- Misconception: Your code trains the models. Replit’s business agreements with model providers cover data handling, and enterprise and paid tiers include stronger privacy terms. Always read the current policy for your plan rather than assuming.
- Misconception: A more powerful model always produces better apps. Context quality, prompt clarity, and project structure often matter more than raw model strength. A messy 200-file project confuses every model equally.
- Misconception: Replit only works for toy projects. Teams ship real internal tools, dashboards, and customer-facing apps on Replit. The limits show up in very large legacy codebases, not in project ambition.
One more worth calling out: people assume that because Replit uses the same underlying models as competitors, all AI coding tools produce identical results. They don’t. Two products using the exact same Claude model can produce wildly different output depending on system prompts, retrieval strategy, tool design, and verification loops. The model is the engine, but the car around it determines how it drives.
How Replit Compares to Other AI Coding Platforms
Almost every major AI coding tool relies on the same small pool of frontier models. What separates them is philosophy. Here’s how the landscape breaks down.
| Platform | Primary Models Used | Model Choice for Users | Core Strength |
|---|---|---|---|
| Replit | Claude family, GPT, Gemini | Abstracted, effort levels instead | Full cloud environment plus hosting and databases |
| Cursor | Claude, GPT, Gemini, custom models | Explicit model picker | Deep local codebase editing for developers |
| GitHub Copilot | OpenAI, Claude, Gemini options | Model selector in chat | Tight IDE and GitHub integration |
| Lovable | Claude and related frontier models | Mostly abstracted | Design-forward web app generation |
| Bolt | Claude family | Limited selection | Fast browser-based prototyping |
| Claude Code | Claude only | Claude model tiers | Terminal-native agentic coding |
| v0 | Mixed frontier models | Limited | React and UI component generation |
The key trade-off is control versus simplicity. Cursor gives developers a dropdown and expects them to know that a reasoning model handles refactoring better than a speed-optimized one. Replit hides that decision because its audience includes founders, marketers, and students who have never opened a terminal. Replit’s bet is that most users would rather say “make it work” than research benchmark scores.
Replit’s other differentiator is the runtime. Cursor edits files on your machine and stops there. Replit runs your app, provisions a Postgres database, stores secrets, handles authentication, and publishes to a live URL. If your goal is a working product rather than a working repository, that end-to-end coverage matters more than model choice.
Picture a solo founder validating a booking app for a small gym. On Cursor, she’d still need to set up hosting, a database, and a domain. On Replit, she describes the app, Agent builds it, and a shareable link exists within an hour. Same underlying models, dramatically different time to first customer.
Getting Better Results From Replit’s Models
Since you can’t hand-pick the model in most cases, your leverage comes from how you prompt and how you structure your project. These habits produce noticeably better output.
Write Prompts the Agent Can Verify
Vague requests produce vague apps. Instead of “build me a CRM,” describe the exact screens, the data fields, and the user flow. Say what a successful result looks like so the Agent’s self-testing loop has something concrete to check against.
Work in Small Increments
Long autonomous runs are impressive, but they compound errors. Ask for one feature at a time, confirm it works, then move on. This also keeps your effort-based costs predictable.
Practical Tips That Consistently Help
- Start with a clear project description, including your preferred stack, before the first build.
- Use Replit’s design or planning mode to lock in structure before generating code.
- Paste exact error messages instead of describing errors in your own words.
- Keep secrets in the Secrets manager so the Agent never hardcodes API keys.
- Commit or checkpoint after every working feature so you can roll back cheaply.
- Switch to Assistant for small fixes rather than waking the full Agent.
- Break large refactors into a written plan first, then execute the plan step by step.
- Tell the Agent explicitly when you want it to test in the browser before declaring success.
- Ask for a summary of changes after big builds so you understand your own codebase.
- Review generated authentication and payment code manually, always.
One more overlooked tip: clean up your project as you go. Every abandoned file, unused package, and leftover experiment eats context space and increases the chance the model retrieves the wrong file. Developers who treat their Replit project like a real codebase get dramatically better AI results than those who let it sprawl.
Costs, Limits, and What Model Choice Means for Your Wallet
Model selection directly drives what you pay. Frontier models cost meaningfully more per token than lightweight ones, and agentic loops consume tokens in bulk because every tool result feeds back into context.
Replit moved from a simple per-request checkpoint model toward effort-based pricing precisely because different tasks consume wildly different resources. A one-line CSS fix and a three-hour autonomous build should not cost the same. Under effort-based billing, a quick edit might cost a few cents of usage while a complex multi-hour Agent 3 session can run into dollars.
Here’s how to think about managing that:
- Use the lowest effort setting that still gets the job done, and escalate only when the task genuinely needs deeper reasoning.
- Batch related requests into one clear prompt instead of firing off ten vague ones that each trigger a full planning cycle.
- Fix trivial things yourself in the editor. Typing a variable name takes two seconds and costs nothing.
- Watch for retry loops. If the Agent keeps failing on the same error, stop it, read the log yourself, and give it the missing information.
- Track usage in your account dashboard weekly so surprises don’t pile up at billing time.
Here’s a realistic scenario. A user builds a simple internal dashboard with authentication and a database. The initial Agent build might consume a few dollars of effort. Then they spend an afternoon making thirty small refinements. If every one of those runs through the full Agent, costs multiply fast. If they use Assistant for the small stuff and reserve Agent for structural changes, the same afternoon costs a fraction as much. The models are identical in quality; the routing choice is what changes the bill.
Where Replit’s Model Strategy Is Heading
Several trends will shape what powers Replit over the next couple of years, and they’re worth watching if you depend on the platform.
First, expect continued multi-model routing to deepen. As the gap narrows between providers, Replit gains negotiating power and reliability by spreading load. Routing intelligence, deciding which model handles which subtask, becomes a competitive advantage in itself. Companies that route well deliver better results per dollar than companies that always reach for the biggest model.
Second, verification will get more attention than generation. The hard problem is no longer writing plausible code; it’s confirming that code actually works. Replit’s investment in browser-based testing, log reading, and automatic error repair points in this direction. Future versions will likely run more sophisticated test suites and security scans automatically.
Third, context management keeps improving. Longer context windows help, but smarter retrieval helps more. Expect better code indexing, dependency graph awareness, and memory of your past decisions so the Agent stops reintroducing bugs you already fixed.
Questions People Ask Most Often
- Can I choose the exact model in Replit? Generally no. Replit abstracts model selection and gives you effort or mode controls instead, though offerings change over time.
- Does Replit use ChatGPT? It uses OpenAI models through the API for certain tasks, but the primary Agent engine has been Anthropic’s Claude family.
- Is Replit’s AI free? There is limited free access, but serious Agent usage requires a paid plan with usage-based billing.
- Can Replit Agent work with existing GitHub repos? Yes, you can import repositories, though very large codebases stretch the limits of any AI agent.
- Which model is best for debugging in Replit? You don’t pick it directly, but higher effort settings route to stronger reasoning models that handle tricky bugs better.
- Will Replit ever train its own frontier model? Unlikely in the near term. The economics favor buying frontier capability and competing on product experience.
The bigger picture is that model identity matters less each year. As frontier models converge in capability, the winners will be platforms that surround those models with the best environment, feedback loops, and safety rails. Replit has bet its entire product on that thesis.
Bringing It All Together
So when someone asks what LLM does Replit use, the honest answer is a layered one: Anthropic’s Claude models do most of the heavy lifting inside Replit Agent, with OpenAI’s GPT models and Google’s Gemini models filling supporting roles depending on the task, cost, and context length. Replit abandoned frontier model training after its early replit-code experiments and now focuses entirely on the orchestration layer, the sandboxed runtime, the deployment pipeline, and the verification loops that turn raw model output into apps that actually run. That decision explains why Replit improves so quickly whenever a new Claude or GPT release lands.
Understanding this stack changes how you use the platform. You’ll write clearer prompts because you know the model needs verifiable goals. You’ll reach for Assistant instead of Agent on small fixes because you understand the cost difference. You’ll keep your project tidy because you know context quality drives output quality. And you’ll stop chasing the myth that one magic model solves everything. The tools will keep getting better, the models will keep converging, and the people who learn to work with these systems thoughtfully will keep shipping faster than everyone else. Start small, iterate often, and let the AI handle the parts you’d rather not type yourself.