TECHNICAL DEEP DIVE

CoreWeave Forge for Product Managers: The Unified AI Development Platform Explained

By Institute of AI PM·15 min read·Oct 6, 2026

TL;DR

CoreWeave launched Forge on September 30, 2026 at Fully Connected in San Francisco. It is a unified AI development layer that connects the five stages of the improvement loop: Run, Observe, Curate, Improve, Evaluate. Forge combines Weights & Biases Models, OpenPipe post-training tooling, and marimo notebooks into one environment on top of CoreWeave compute. Early adopters include MasterClass and Canva. The free tier is live; Pro starts at $60 per month. For AI PMs managing teams that iterate on models or agents in production, this is the most opinionated all-in-one alternative to stitching together five separate tools.

The AI PM Minute

One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.

No fluff. Unsubscribe anytime.

What CoreWeave Forge Is and Why It Launched Now

CoreWeave started as a GPU cloud provider: companies rented H100 clusters to train and serve models. Forge is CoreWeave's move up the stack. Instead of just renting compute, CoreWeave now sells a development environment that sits on top of that compute and makes the full iteration cycle faster.

The timing is not accidental. Two forces converged in 2026. First, training and fine-tuning moved from research teams to product teams. PMs at companies like Canva and MasterClass are now responsible for the quality of models that ship in production, not just the features those models power. Second, the tool landscape became fragmented. Teams use W&B for experiment tracking, a separate inference provider for serving, a third tool for data curation, another for evals, and often a notebook environment that is disconnected from all of them. Every handoff between tools introduces latency and context loss.

Forge positions itself as the answer: one platform where the output of a production run (the observed failures, the flagged outputs, the curated examples) flows directly back into the next training or fine-tuning job, without any export-import friction.

The headline from the launch

"Every production run should make your model better. CoreWeave Forge turns the AI loop from a one-time deployment into a compounding improvement system." This is the core product thesis: production is not the end of development, it is the source of data for the next iteration.

The Five Stages of the Forge Loop

Forge is structured around five stages that form a closed loop. Understanding each stage helps you evaluate whether Forge actually solves your team's bottleneck or whether you already have a good solution for most of these stages and only need one or two.

Run

What it does: Inference and agent execution on CoreWeave compute. Forge captures every request and response, including latency, token counts, and metadata. This is the source layer for everything else in the loop.

PM implication: If you are already running on a different inference provider, this is the biggest switching cost. Run requires CoreWeave compute. The rest of Forge can integrate with external run environments, but you get the tightest loop on CoreWeave infrastructure.

Observe

What it does: Real-time monitoring of model behavior in production. Powered by Weights & Biases tracing. You see token usage, latency distributions, error rates, and output quality signals across every inference call.

PM implication: If you are already on W&B for experiment tracking, Forge's Observe layer will feel familiar. The differentiator is that it is connected to your production runs automatically, not a separate integration you have to build.

Curate

What it does: Surfacing and labeling the most useful examples from production runs. Forge uses embedding similarity and model-graded scoring to identify which outputs represent failure modes, edge cases, or high-value examples worth adding to your training set.

PM implication: This is where Forge is most differentiated. Most teams curate training data manually or with custom scripts. Forge makes curation a first-class workflow: filter by quality score, tag by category, send to a labeling queue, and push directly into a fine-tuning job.

Improve

What it does: Fine-tuning and post-training workflows powered by OpenPipe. Forge integrates OpenPipe's structured output fine-tuning tooling and marimo notebooks for exploratory analysis. A curated dataset becomes a fine-tuning job in a few clicks.

PM implication: OpenPipe specialized in fine-tuning models on structured outputs, which matters a lot for product use cases where you need consistent JSON or a specific response format. This is not general pre-training infrastructure. It is designed for the fine-tuning and post-training loop that product teams actually run.

Evaluate

What it does: Running evals on new model versions before they graduate to production. Forge supports LLM-as-judge eval frameworks and integrates with benchmark suites. Results are visible in the same dashboard as production monitoring, so you can compare a candidate model against the current production model on the same distribution of real inputs.

PM implication: The eval-in-production-context story is compelling. Running evals on synthetic benchmarks often misses the actual failure distribution you see in production. Forge lets you build evals from real production examples, which increases the chance your eval scores predict real-world quality.

How Forge Compares to the Existing Tool Stack

Before Forge, a typical production AI team assembled tools from different vendors for each stage of the development loop. Forge does not replace every tool in that stack, but it changes the build-vs-buy calculus for teams that are assembling a stack from scratch in 2026.

vs. Weights & Biases standalone

W&B now ships as a component inside Forge (W&B Models). If you are a current W&B user, Forge adds inference infrastructure, curation, and OpenPipe post-training on top of W&B's experiment tracking you already use. The integration is native, not a connector. The question is whether you want to move your compute to CoreWeave.

vs. MLflow

MLflow is open source and compute-agnostic. Forge is opinionated and CoreWeave-first. MLflow wins if you need portability across clouds or on-premises. Forge wins if you want a managed experience and your compute is already on CoreWeave or you are willing to move it.

vs. Comet ML

Comet focuses on experiment tracking and model monitoring. It does not have the curation or post-training components Forge offers. For teams that need end-to-end, Forge covers more surface area. For teams that only need monitoring and tracking, Comet is lighter.

vs. custom scripts and S3

Most product teams at the stage of running fine-tuning cycles are doing data curation and job management with custom Python scripts and object storage. Forge replaces this with a product. The question is whether the managed workflow is worth the platform dependency.

Learn to Evaluate AI Infrastructure as a PM

The AI PM Masterclass covers how to assess AI infrastructure decisions, evaluate vendor trade-offs, and translate technical choices into product outcomes. Taught live by a Salesforce Sr. Director PM.

Pricing and What Each Tier Actually Gives You

Forge launched with three tiers. The pricing is straightforward, but the meaningful limits are in the free tier, which should tell you whether Forge is worth evaluating at your stage.

Free

$0

Access to the Forge dashboard, experiment tracking via W&B integration, and limited inference runs. Designed for individual developers and small teams evaluating the platform. Data retention and concurrent job limits apply.

Useful for evaluation. Not sufficient for a production workload running more than a handful of fine-tuning jobs per month.

Pro

$60 per month per seat

Full access to all five stages of the loop. Higher data retention, more concurrent jobs, priority support, and access to OpenPipe fine-tuning infrastructure. This is the tier MasterClass and Canva likely started on before moving to Enterprise.

The right tier for teams that have at least one model or agent running in production and are beginning to iterate on it systematically.

Enterprise

Custom

Custom compute contracts, SLA guarantees, SSO, dedicated support, and extended data retention. Designed for teams running at scale on CoreWeave infrastructure who need contractual commitments.

Appropriate once you are running enough fine-tuning volume that the per-seat pricing would be more expensive than a custom contract, or when your security team requires SSO and audit logging.

Five PM Questions Before Adding Forge to Your Roadmap

Forge is a genuine addition to the AI infrastructure market, but it is not the right choice for every team. Here are the five questions to answer before you commit.

1

Is your compute already on CoreWeave, or are you willing to move it?

The Run stage ties Forge to CoreWeave infrastructure. If you are on AWS, GCP, or Azure and have no plans to move, you lose the tightest part of the loop. The curation and eval tools can integrate with external inference, but the observability layer is most powerful when CoreWeave captures every inference call natively.

2

Are you running fine-tuning or post-training cycles in production?

Forge is designed for teams that iterate on models, not just teams that call an API. If you are using a hosted model via an API and have no fine-tuning in your roadmap, Forge's Improve stage adds no value. The observability and eval stages are still useful, but you do not need a platform this opinionated just for monitoring.

3

How big is your data curation bottleneck?

The Curate stage is where Forge has the strongest differentiation. If your team is spending significant PM or engineering time on data labeling workflows and is losing quality signal because you lack a systematic curation process, this is the stage most likely to deliver ROI.

4

What is your current W&B relationship?

If you are already paying for W&B Teams or W&B Enterprise, there is a pricing conversation to have. Forge includes W&B Models, so the incremental cost of Forge over your existing W&B spend depends on your current contract. Talk to both CoreWeave and W&B before signing anything.

5

What is your exit cost if Forge does not work out?

Evaluate your data portability before you start. Can you export your curated datasets, eval results, and model checkpoints in a standard format? Can you reproduce your fine-tuning jobs on a different infrastructure provider? Platforms that make ingress easy and egress hard are a vendor lock-in risk, especially if CoreWeave changes pricing or gets acquired.

What Forge Tells Us About the Direction of AI Infrastructure

CoreWeave's move is a signal about where the AI infrastructure market is heading. Compute providers that only sell GPUs are commodity businesses. Margins compress as competition increases and customers negotiate harder. Moving up the stack into developer tooling creates stickier relationships and higher margins.

The acquisitions that built Forge tell the story: Weights & Biases brought the experiment tracking and observability layer; OpenPipe brought post-training expertise specifically for structured outputs; marimo brought an open-source notebook environment with a growing developer community. Each acquisition addressed a different stage of the improvement loop, and Forge is the product that connects them.

The broader trend: vertical integration in AI infra

AWS, Google, and Azure have all moved from raw compute to managed training and inference services. CoreWeave is following the same pattern. The question for AI PMs is whether the vertical integration creates enough value to justify the reduced portability.

What early adopter choices signal

MasterClass and Canva are both consumer-facing product companies, not AI labs. Their adoption suggests Forge is targeting product teams that do fine-tuning as a means to better product quality, not research teams building foundation models. That is a much larger market.

The open-source hedge

Forge includes marimo, an open-source notebook. This is a customer acquisition strategy as much as a product decision: open-source components reduce adoption friction and create a path from the free tier to the paid platform. Watch how aggressively CoreWeave invests in marimo's community.

Competition incoming

AWS SageMaker, Google Vertex AI, and Azure AI Studio all cover overlapping ground. The difference is that CoreWeave is not a hyperscaler with 50 other products competing for internal engineering attention. Focused infrastructure providers often ship faster on their core use case. That advantage is real but not permanent.

Build AI Products That Actually Improve in Production

The AI PM Masterclass teaches you how to evaluate AI infrastructure, design feedback loops, and make build-versus-buy decisions with confidence. Taught live by a former Apple and Salesforce Sr. Director PM.

Before you go: get the AI PM Minute

One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.

No fluff. Unsubscribe anytime.