Prompt-First Product Development: How AI PMs Ship Features Through Natural Language
TL;DR
Traditional PRDs describe what the software should do. In prompt-first development, the system prompt or agent instruction IS the feature. You are not writing a spec that engineers then implement: you are writing the implementation directly in natural language. This shift changes how PMs write requirements, how they collaborate with engineers, what "done" looks like, and where the failure modes are. This guide explains the methodology, where it works, and the three ways teams get it wrong.
The AI PM Minute
One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.
No fluff. Unsubscribe anytime.
Why Traditional Feature Specs Break Down for AI Features
A traditional PRD describes desired system behavior, which engineers then translate into code. The gap between the spec and the code is the engineer's job. For software features, this gap is managed through well-understood processes: acceptance criteria, automated tests, code review.
For AI features, the gap disappears in a specific direction: the model's behavior is largely determined by what you tell it in the system prompt, the tool definitions, and the few-shot examples you provide. The prompt is not documentation for the feature. The prompt is the feature. When you change the prompt, you change what the feature does, without a single line of code changing.
This creates a structural mismatch with traditional PM workflows. Writing a PRD for an AI feature and then handing it to an engineer to "implement" conflates two things: the application scaffolding (API calls, UI, data pipeline) and the actual feature behavior (the prompt). Engineers can build the scaffolding. Only someone who understands the desired behavior in precise, testable natural language can write the prompt. That person is increasingly the PM.
What engineers own in AI features
API integration, latency optimization, cost management, tool definitions, retrieval pipeline, caching, fallback logic, monitoring, evaluation harness, deployment. This is real engineering work that deserves careful specs.
What PMs own in AI features
The system prompt, the output format, the persona, the guardrails, the few-shot examples, the task decomposition logic, the escalation conditions. This is product behavior, specified in natural language.
Where the old model fails
A PM writes acceptance criteria. The engineer writes a system prompt to satisfy them. The PM tests it and says 'it sounds off.' Neither party had a shared language for what the prompt was supposed to accomplish. Iteration takes weeks.
What prompt-first fixes
The PM writes the prompt as the primary artifact, versioned and reviewed like code. The acceptance criteria become: does the model behave as this prompt specifies? The engineer builds the infrastructure to deploy and measure it.
The Prompt-First Framework: Writing Prompts as Product Specs
A prompt-first feature spec has five components. Each maps to a section of the eventual system prompt, but it is written in a format that can be reviewed, versioned, and tested before the application scaffolding exists.
1. Role and context
Who is the model acting as, and what does it know about the product context? This is not a persona for its own sake. It sets the baseline of knowledge the model should assume it has and the perspective it should reason from.
You are a customer support agent for Acme CRM. You have access to the user's account history, open tickets, and product documentation. You do not have access to billing records or the engineering backlog.
2. Task definition
What is the model being asked to do, stated precisely enough that you could write a test for it? The test requirement is the discipline: if you cannot write a test case from this definition, it is not precise enough.
Given a user question and their account history, provide a resolution path. If the issue is resolvable from documentation, provide the steps. If it requires human review, classify it by tier (billing, technical, account) and summarize the issue in one sentence.
3. Output format and constraints
How should the response be structured? What must it include? What must it never include? Be explicit about format, length, tone, and any content constraints. This section prevents the model from being unhelpfully verbose or dangerously terse.
Respond in plain English, under 150 words. Do not include legal disclaimers. Do not promise resolution timelines. Always end with a confirmation question: 'Does this answer your question?'
4. Edge cases and refusals
What should the model do when the user asks something outside scope? Explicit refusal conditions are the safety rails of the feature. Teams that skip this section discover the edge cases in production.
If the user asks about competitor products, redirect to Acme features without disparaging competitors. If the user asks for a refund, acknowledge the request and transfer to billing without making commitments.
5. Examples (few-shot)
Two to five examples of ideal input/output pairs. These are the most high-leverage lines in the prompt. They calibrate the model's behavior better than any amount of instruction prose. Write them for the hard cases, not the easy ones.
Provide one example of a password reset question handled well, one example of a billing question handled with a clean handoff, and one example of an out-of-scope question redirected gracefully.
How Prompt-First Changes the PM-Engineering Workflow
The biggest workflow change is what gets reviewed in design review and what gets reviewed in code review. In a prompt-first team, the prompt spec is a first-class artifact that goes through design review before any code is written. This is not documentation-for-its-own-sake. It is the equivalent of reviewing a UI mockup before building the component.
PM writes the prompt spec
Using the five-component framework above. This is the primary deliverable. It takes roughly the same time as writing a solid PRD section, but produces something executable rather than something that needs to be interpreted.
PM and engineer review the prompt spec together
The engineer reads it and identifies: what tooling is needed to evaluate whether this spec is met, what infrastructure is needed to deploy it, what edge cases are not covered. Both parties reason from the same artifact.
PM runs the prompt against the model directly
Before the engineer writes any scaffolding, the PM runs the spec as a system prompt and tests it against representative inputs. This is a 30-minute prototype step. If the prompt does not work in a chat interface, it will not work in the product.
Engineer builds the scaffolding to the spec
API integration, retrieval pipeline, caching, monitoring. The engineer's job is to make the prompt run at production scale with the right latency, cost, and reliability properties.
Evaluation is written from the spec
The acceptance criteria are derived directly from the prompt spec's task definition and constraints. Eval test cases are written by the PM alongside the spec, not after the engineer ships. This makes the feedback loop tight.
Learn to Ship AI Features That Actually Work
The AI PM Masterclass teaches prompt-first product development, evaluation design, and the full stack of AI PM skills. Taught live by a Salesforce Sr. Director PM.
The Three Ways Teams Get Prompt-First Wrong
Mistake 1: Treating prompts as an implementation detail
Symptom: The PM writes the PRD, the engineer writes the prompt, and the PM only sees the prompt when reviewing the demo. The feature ships with behavior that no one deliberately chose.
Fix: The prompt spec is a design review artifact. The PM writes a draft prompt before the engineer writes any code. If the PM has never touched the prompt, they have not finished their part of the work.
Mistake 2: Writing prompt specs without testable acceptance criteria
Symptom: The spec says 'be helpful and accurate.' No one can write an evaluation for 'helpful and accurate.' The feature ships, users complain it is not helpful or accurate, and no one knows how to fix it systematically.
Fix: Every clause in the task definition should have an associated test case before the sprint starts. 'Resolve password reset issues from documentation' means you write a password reset test case before the engineer builds anything.
Mistake 3: Not versioning prompts like code
Symptom: Someone edits the system prompt in production to fix a complaint. Three weeks later, a different behavior regression appears. No one remembers what changed. The prompt is not in version control. The root cause investigation is a manual archaeology project.
Fix: Prompts live in the repository alongside code. Prompt changes go through the same review process as code changes. The system prompt is never edited in a provider's playground and copy-pasted to production. It is a file, versioned, reviewed, and deployed.
Where Prompt-First Works and Where It Does Not
Prompt-first is a strong fit for a specific class of AI features. It is not a universal development methodology.
Strong fit
- +Customer-facing conversational features (support, onboarding, copilots)
- +Document summarization and extraction pipelines
- +Content generation with constrained formats (summaries, descriptions, reports)
- +Classification and routing logic with explicit rules
- +Single-turn or short multi-turn interactions with clear task scope
Weak fit
- -Multi-step agentic workflows with complex state management
- -Features where the logic is primarily algorithmic (recommendation engines, pricing models, search ranking)
- -Real-time processing at latency budgets under 200ms
- -Features requiring guaranteed determinism (financial calculations, legal compliance checks)
- -Long-horizon planning tasks where the model must remember state across many turns
Putting It Together: A Prompt-First Feature Sprint
Here is what a two-week prompt-first feature sprint looks like in practice, using a customer-facing email draft generator as an example.
PM writes the prompt spec
Role, task, output format, edge cases, and three to five few-shot examples. Uses the chat interface of the target model to validate that the spec produces the desired behavior before sharing with the team.
Prompt spec design review
PM presents the prompt spec and live model demos to the team. Engineer identifies infrastructure needs. Designer checks output format against the UI. Spec is updated based on feedback.
Eval test cases written
PM writes 15 to 20 test input/output pairs covering the core use case, edge cases, and refusal conditions. These become the acceptance criteria. They live in the repo.
Engineer builds scaffolding
API integration, UI, caching, monitoring, evaluation harness that runs the test cases. PM continues refining the prompt in parallel as they learn from more testing.
Eval run and prompt refinement
Evaluation harness runs all test cases against the current prompt. PM and engineer review failures together. Prompt is revised until pass rate meets the agreed threshold.
Staged rollout
Feature ships to 5% of users with logging enabled. PM monitors the prompt version, model, and qualitative feedback. The prompt spec version is tagged in the release.
Master the Full AI PM Toolkit
Prompt-first development is one of a dozen methodologies the AI PM Masterclass teaches hands-on. Stop learning AI PM theory. Start shipping AI products.
Related Articles
Before you go: get the AI PM Minute
One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.
No fluff. Unsubscribe anytime.