Icon Rounded Closed - BRIX Templates

The Model Is Not the Whole Product

The Model Is Not the Whole Product

By Vishwas Lele
Co-Founder & CEO, pWin.ai (WordX) | Board Member, Applied Information Sciences | Microsoft Regional Director


What Brendan Foody gets right about AI applications, and what proposal teams should expect beyond a capable model.


I have been thinking about Brendan Foody’s argument that “the model is the product.”

In his 20VC interview with Harry Stebbings, the Mercor CEO challenges application builders to explain where lasting value will come from. He places particular weight on infrastructure, network effects, and working closely with enterprises to understand their workflows and evaluate AI against their actual work.

As a co-founder of pWin.ai, I have a stake in this debate. We help federal contractors bring requirements, capture strategy, and supporting evidence into proposal development. That makes the question practical for us: what should a specialized application deliver beyond access to a capable model?

I think Foody is right about a risk that application builders should take seriously. I am less convinced that it leaves little lasting value for specialized applications.

For proposal teams, I believe that value comes from helping them keep track of the work: which requirements apply, what evidence supports the claims, what the team has decided, and what needs reconsideration when something changes.

Generating the words is part of the job. Helping an organization develop, review, and stand behind its response is the larger responsibility.


Where I think Foody is right


Suppose a product exists primarily because today's models cannot reliably perform a particular task. The company builds prompts, routing logic, and workarounds to close that gap.

Then a new model arrives and handles the task directly.

The engineering may have been difficult. Customers may have received real value. But neither guarantees that the advantage will last. Calling the implementation an "agent architecture" does not change that possibility.

Sequoia's Julien Bek makes a related argument in Services: The New Software: selling completed work can allow a business to benefit from better models because those models reduce the cost of delivering the outcome. That does not settle which business model will succeed, but it is a useful way to think about what customers are paying for.

I also find Foody's emphasis on working deeply inside customer organizations persuasive. A demonstration may not reveal an approval that happens outside the documented process, an exception everyone knows about but nobody has written down, or why an apparently acceptable answer would never be used.

The opportunity is to turn that understanding into repeatable product behavior. Otherwise, we may have created a valuable consulting engagement without creating a durable software business.

And when a better model lets us remove a complicated part of our implementation, we should welcome that. Preserving our code is not the objective.

A proposal makes the distinction easier to see

‍
Consider an illustrative solicitation with a 25-page technical response limit.

A contractor has an ambitious modernization approach, but the customer is especially concerned about continuity during transition. The contractor also has stronger evidence of successful transitions than of delivering that modernization approach at comparable scale.

Where should the next half-page go?

Answering that requires weighing the evaluation criteria, the available evidence, the proposed solution, and what to remove elsewhere. Sometimes the better choice is a less exciting claim that the team can substantiate.

The resulting outline embodies a strategy.

Now suppose an amendment changes a requirement. During review, someone also discovers that a past-performance example attributes an entire program to the contractor when its team delivered only one workstream.

The team must determine which arguments remain valid, which passages need changing, and who must approve the revised claims.

A capable model can help with all of this. I would not argue that these decisions are inherently beyond AI.

But the work still needs continuity, context, and experience. A good response depends on a deep understanding of why the customer issued the solicitation: what problem it is trying to solve, what constraints it faces, and what concerns shaped its requirements.

Some of that understanding comes from customer conversations, capture work, and experience with earlier programs. It is not always stated in the solicitation. The team needs to carry that context, and the reasoning it informs, through planning, drafting, and review.

The team also needs to know which requirement is current, which evidence is appropriate, what has been approved, and whether a later change has affected an earlier decision.

This is what I mean by maintaining the "state of the work." It is not simply saving the latest document. It is keeping the requirements, customer context, evidence, and decisions behind that document understandable as the proposal develops.

A useful application should help the team answer three questions: why does this argument belong in this proposal, what supports it, and what needs reconsideration when something changes?

The application I am describing therefore needs to do more than provide a thin layer over a model. It should connect the model's capabilities to the customer context, evidence, and decisions accumulated through the proposal process, and help the team keep them current as the work changes.

What about SharePoint and Copilot?

For an organization using SharePoint and Microsoft 365 Copilot, the practical question may be: why add a specialized proposal application?

That is a fair question.

Microsoft's tools should not be reduced to a chatbot with a few uploaded files. With appropriate licensing and configuration, Copilot can work with organizational content. SharePoint agents can answer questions using material the user is authorized to access, and Copilot Studio supports agents that combine knowledge, tools, and actions.

The distinction I would draw is between having those capabilities available and having them organized around the proposal process.

Access to a past-performance document does not, by itself, establish that it supports a particular claim. Access to a solicitation does not settle how the team should allocate limited pages. A record of an earlier decision does not establish that the decision remains appropriate after an amendment.

These are requirements for the overall solution, not reasons to claim that Microsoft—or another platform—cannot address them.

For pWin, the responsibility is to bring proposal-specific planning, evidence, drafting, and review together in a way that reduces the work customers must assemble and maintain themselves. The question is not which tool can produce the most impressive paragraph. It is which approach helps the team reach a response it can approve and use, with less total effort.


What pWin does today


Today, pWin's Content Plan gives teams a place to establish customer priorities, win themes, solution choices, and writing direction before drafting. The Knowledge Repository brings relevant organizational material into that work and provides source references for generated content.

Our annotated outlines map solicitation requirements to planned response sections. Draft generation follows the team's plan, while Compliance, Citation, and Hallucination Reports help reviewers examine requirement coverage, supporting sources, and potentially unsupported claims. These are existing capabilities that support expert review; the team retains responsibility for the final response.

A concrete example comes from our published PROSOFT case study. PROSOFT reported saving up to five days in moving from a blank page to a working first draft. Its team highlighted the combination of the Knowledge Repository, Outline Builder, Content Plan, and draft generation—not simply the speed of generating text. The case study also describes the team being able to pursue opportunities it previously had to pass on.

That is one customer's reported experience, not a promise that every team will achieve the same result. But it illustrates the kind of value I mean: less time assembling the work and more capacity to develop a response.

What the research does—and does not—show

Harvey's September research offers a useful example of how the system around a model can affect its performance.

For a mergers-and-acquisitions diligence task, Harvey compared a conventional setup with a purpose-built system that divided document review among agents and assembled their findings. Across seven coordinating models, the average share of evaluation criteria satisfied increased from 23.3% to 62.4% on 50 held-out synthetic data rooms. An AI judge scored the results against expert-defined criteria.

These are company-reported benchmark results, not transaction success rates. The more elaborate setup also cost more to run for most of the models tested. The conclusion I draw is that application design can materially affect performance, while quality and cost still need to be considered together.

OpenAI's September 17 announcement of Astra for Law illustrates another side of the debate. It combines a model with legal search, specialized instructions, tools, and professional controls. The announcement also says that Harvey and Legora will be able to build on that foundation.

My reading is that a foundation-model company can both compete with specialists and enable them. Better models and better applications are not mutually exclusive.

Evidence that application design improves a result does not, by itself, establish a lasting business advantage. But it does show why evaluating the model alone is not enough.

What Mercor teaches us about expert judgment

It would be easy to respond to Foody by saying that professional work requires judgment, and therefore specialized applications and human experts will always be protected.

Mercor's work should make us more careful with that argument.

In September, Mercor and the SkyRL team reported improving model performance through additional training on expert-created tasks in management consulting, investment banking, and corporate law. The training environments and prompts were separate from those used in the benchmark. These results concern other professions, but they demonstrate how expert-defined work can contribute to model improvement.

Mercor also emphasizes that useful evaluations need input from people who actually do the work. Engineers alone may not know which omissions, exceptions, or mistakes matter most.

The lesson I take is not that every application company needs to train its own model. It is that experienced practitioners should help define what good work looks like and how to evaluate it.

Here is a thought experiment about proposal development.

Imagine 100 experienced federal-proposal teams helping create realistic exercises using fictional contractor records, solicitations, amendments, and capture notes. The task would be to prepare a supported response, identify missing information, and incorporate a change.

Some judgments would be relatively concrete. Did the response use a superseded requirement? Does the staffing table contradict the narrative? Does the evidence support the claim?

Other judgments would allow several good answers. Two different page allocations might both be defensible. Experts would need to explain the tradeoffs rather than grade against one preferred response.

Some questions could not be answered from the supplied information. Has the subcontractor committed the proposed staff? Has leadership accepted the delivery risk? Recognizing the gap and asking the right question should count as good performance.

The difficult work is agreeing on what good performance looks like without rewarding the wrong behavior. Counting citations could encourage irrelevant citations. Counting requirements mentioned could reward repetition rather than a credible approach.

These assessments cannot rely on heuristics alone. Simple rules can help identify issues, but judging the quality of a response requires weighing the customer's priorities, the evidence, and relevant experience. A citation can point to a relevant document without supporting the specific claim. A requirement can be mentioned without explaining how the contractor will meet it. Expert judgment is needed to define and evaluate those distinctions, including when a capable model contributes to the work.

This is why I value our ongoing work with Shipley Associates: We Help Companies Win Business. We already apply that expertise to product behavior, review criteria, and how customers use pWin in their proposal process. The work is about putting practitioners' judgment into practice, not simply attaching a methodology's name to generated text.

We should be open to how much AI can learn while remaining clear about who has authority to approve the work. Those are different questions.

Does the tenth proposal benefit from the previous nine?

For every customer, after each draft, we ask ourselves a question:

Are we helping this organization make its next proposal easier and better because of what it learned from this one?

Over time, that becomes a larger question: does the tenth proposal benefit from the previous nine?

We are not simply asking whether pWin released a better version, whether the underlying model improved, or whether users became more familiar with the interface. Has the organization retained useful, approved lessons? Can it reuse stronger evidence, avoid repeating the same corrections, and understand the reasoning behind earlier decisions?


Beyond today's planning, drafting, and review capabilities, I see this as an important standard for the next stage of proposal software.

Consider a reviewer who changes:

"We delivered the enterprise migration."

to:

"We led the migration workstream within the prime contractor's broader program."

The valuable lesson is not the revised sentence alone. It is the accurate scope of the team's contribution.

That distinction should not need to be rediscovered on the next bid. It should remain supported by evidence and clear enough for another team to use appropriately.

The same applies to other lessons. An old customer insight should remain dated and attributed. A project-specific decision should not quietly become a company-wide rule. A preference for shorter paragraphs should not be confused with a correction to a material fact.

Retaining approved organizational knowledge is different from training a model on customer documents or edits. pWin does not train or fine-tune models on customer content.

The broader goal is to make learning useful within the planning, knowledge, and review experiences teams already use. Important corrections should help the organization recognize and avoid the same problem again, rather than disappear when the proposal is submitted.

Supporting the full workflow can be valuable. For me, the measure is how much it helps teams use reliable knowledge, make better decisions, and carry what they learn into the next pursuit.

What customers should measure

A fast first draft is only part of the answer.

I would evaluate the full path to a response the team is ready to approve. How many material requirements were missed? How many unsupported claims survived? How much expert review was required? How difficult was it to incorporate an amendment or a change in strategy?

The comparison should include the effort required to prepare the information, configure the tools, and verify the output. It should account for total cost and elapsed time—not just generation time.

Use representative solicitations, including ones the system was not tuned around. Have experienced reviewers judge the results. Repeat the tests, because one good run is not enough.

Count false alarms as well as missed issues. A system that flags everything can create considerable work while appearing cautious.

These standards should apply whether the team uses a specialized application, a configured enterprise assistant, or a combination of tools.

For me, the practical measure is whether the technology reduces the total effort required to produce work the organization can stand behind.

What I take away

Foody's argument is useful because it asks application builders to examine what customers are really paying for.

I do not think the answer is that the model absorbs all meaningful value. I also do not think saying "we understand the workflow" is enough.


For pWin, the standard is the customer's work: stronger support for claims, less repeated review, better use of evidence, and decisions that remain understandable as a proposal changes.

Every pursuit should leave the organization with better evidence, clearer decisions, and fewer problems to rediscover next time.

That is the value I believe we should build and measure.

What does your team repeatedly correct in proposals—and how do you keep that correction from getting lost before the next one?

Share this post

Let us Help You Win

Whether you need training or consulting support to facilitate your next win, click one of the options below and let's get started.

Gallery - Elements Webflow Library - BRIX Templates