According to an MIT study, 95% of enterprise generative AI projects fail. Gartner estimates more than 40% of agentic AI projects will be abandoned by the end of 2027. Yet every week, decision-makers sign contracts after spectacular demonstrations without knowing how to distinguish a working solution from polished stagecraft.
The phenomenon has a name: AI washing. The US SEC has begun penalizing companies for false AI capability claims, with fines totaling $400,000 in its first cases in March 2024. The problem extends beyond finance: Gartner estimates that of thousands of identified agentic AI vendors, only about 130 offer real capabilities.
This article gives you 12 practical questions to separate substance from marketing during an AI demo. You do not need technical expertise to use them.
TL;DR — Most AI demos are designed to impress rather than inform. To evaluate objectively, ask about training data, real error rates and production latency, and insist on testing your own data. These 12 questions let you do so without technical skills.
Why Most AI Demos Do Not Reflect Reality
The Gap Between Demo and Production
An AI demo is a showcase: the best possible scenario, carefully selected data and a controlled environment. Production is radically different, with noisy data, edge cases, load increases and third-party service outages.
The gap is not always intentional. Some vendors sincerely believe in their products but have tested only clean datasets. Others are experts at staging. The result for buyers is identical: decisions based on a distorted picture.
In France, only 4% of organizations report significant financial benefits from AI investment, according to Squid Impact's compilation of 2024–2025 industry reports. This reflects a systemic problem in solution selection and evaluation.
The Wizard of Oz Technique
The term comes from 1980s computing, coined by researcher John F. Kelley at Johns Hopkins University. Users interact with what they believe is an automated system while a hidden human operator produces responses manually.
IBM already used this technique in the 1970s to simulate speech recognition: users spoke to a computer believing it transcribed speech, while a person typed in real time in the next room.
Today, the technique has evolved. Invisible human intervention can disguise actual automation during some AI demos. Nate Inc. is an extreme example: its founder raised $42 million claiming AI processed transactions, while manual workers actually made purchases. Claimed automation exceeded 90%; in reality, it was near zero. The SEC and DOJ brought proceedings in April 2025.
AI Washing: Documented and Penalized
AI washing means claiming capabilities the product does not possess. It ranges from marketing embellishment—renaming rules-based systems AI—to outright fraud.
In March 2024, the SEC simultaneously penalized Delphia (USA) Inc. and Global Predictions Inc. for misleading claims about AI in investment processes. In January 2025, Presto Automation Inc. became the first publicly traded company penalized for AI washing: SEC analysis found its supposedly automated speech technology relied on substantial human intervention and belonged to an undisclosed third party.
Gartner introduced agent washing to describe vendors relabeling existing chatbots, RPA and assistants as AI agents without actual agentic capabilities. The phenomenon affects most of the market.
Four Categories of Warning Signs During a Demo
Visual Signals: What You See—or Do Not
Watch what presenters omit as closely as what they show. Look for these warning signs:
| Warning sign | What it may hide | Question to ask |
|---|---|---|
| Always the same dataset | System works only with prepared data | “Can we test our own data?” |
| Presenter avoids edge cases | Model fails on exceptions | “What happens with incomplete or ambiguous input?” |
| Responses seem instantaneous | Overprovisioned demo environment | “What is average production response time?” |
| Polished interface, vague results | Frontend investment rather than model quality | “What is documented accuracy?” |
| No errors in 30 minutes | Scripted scenario | “Show a case where it gets something wrong.” |
Verbal Signals: Language That Should Raise Concern
Certain phrases signal immature or oversold solutions. “Our AI understands,” “our algorithm learns in real time” and “our system is intelligent” are marketing metaphors rather than technical descriptions.
A serious vendor discusses precision, recall, F1 score, 95th-percentile latency and training-data volume. It quantifies limitations rather than denying them.
Commercial Signals: Pressure to Sign
A vendor pushing for a quick signature after a demo without offering a test on your data has something to hide—or does not believe its own product will survive real-world testing.
The classic sequence is an impressive demo, artificial urgency such as next month's price increase, then a multiyear contract without a POC. This should trigger suspicion, not enthusiasm.
Technical Signals: Missing Documentation
A mature AI product has accessible technical documentation: architecture, data sources, performance metrics, error-handling policy and roadmap. If it does not exist or is not yet available, the product is probably not production-ready.
Twelve Questions to Ask During an AI Demo
Questions About the Model and Data: 1–4
Question 1: “What type of model do you use, and is it proprietary or based on a foundation model?”
This establishes whether the vendor developed its own model or built an application layer over GPT-4, Claude or Gemini. Neither is inherently better, but cost, dependency and control implications differ radically. A vendor unable to answer clearly probably does not master its own stack.
Question 2: “What data was the model trained or fine-tuned on?”
Model quality directly depends on training-data quality and representativeness. If the vendor fine-tuned on industry data, ask which data, how much and how it was annotated. Data mismatched to your industry or geography will degrade production performance.
Question 3: “Will our data be used to train or improve your model?”
This concerns intellectual property and GDPR. If your data feeds the vendor's model, it indirectly benefits other customers, including competitors. The answer must be clear and contractual, rather than “in principle, no.”
Question 4: “Is the model explainable or a black box?”

An explainable model lets you understand why a decision was made. For credit scoring, medical diagnosis and regulatory compliance, explainability is not optional: the European AI Act, which entered into force in June 2024, requires it for high-risk systems. If the vendor cannot explain decisions, check whether your use case falls into a regulated category.
Questions About Actual Performance: 5–8
Question 5: “What is your documented accuracy, and which dataset was used to measure it?”
The difference between 80% and 95% accuracy may seem marginal. In practice, the first system is wrong four times as often. Ask which benchmark or test set produced the figure and whether it reflects production conditions or a laboratory environment.
Question 6: “What is 95th-percentile latency in production?”
Average response time is misleading. A system averaging 200 ms but taking 5 seconds one time in twenty harms user experience. P95 reveals what users experience in unfavorable cases; this is the figure that matters for infrastructure sizing.
Question 7: “What happens when the model cannot answer?”
Every model has limits. A well-designed system detects uncertainty and escalates to a person or signals low confidence. A poorly designed system hallucinates, delivering plausible false answers as confidently as correct ones. Failure handling reveals more about maturity than success cases.
Question 8: “Can you show a case where the system fails?”
A vendor refusing to show limitations or claiming its AI never makes mistakes is lying or deceiving itself. Every system has weaknesses. A mature vendor knows, documents and actively manages them. This transparency signals reliability more strongly than any performance metric.
Questions About Integration and Operations: 9–12
Question 9: “How does the system integrate with our existing stack?”
AI operating in a silo has no business value. Ask which APIs and data formats are supported and what authentication and security constraints apply. Check that integration does not require proprietary middleware creating additional dependency.
Question 10: “What is the three-year total cost of ownership, including infrastructure, maintenance and retraining?”
AI costs extend beyond licenses: compute infrastructure—GPUs and cloud—model maintenance as performance drifts, periodic retraining, data annotation and support. Request a detailed breakdown. A vendor unable to provide one probably lacks sufficiently longstanding production customers to know these costs.
Question 11: “What happens if we decide to stop the service?”
Exit options are often overlooked. Can you export data, and in what format? Do you own models fine-tuned on your data? How long is transition? A vendor making exit difficult or impossible builds its business on lock-in rather than product value.
Question 12: “Can we run a 30-day POC on our own data before committing?”
This is decisive. A 30-day proof of concept on real data, with predefined success metrics, is the only reliable way to validate promises. Every serious vendor accepts the principle. Gartner notes that 70% of AI POCs never reach production, demonstrating that the demo-to-reality gap is the norm rather than the exception.
The Evaluation Matrix: Score an AI Demo in 15 Minutes
Five Weighted Criteria
Use this matrix after every demo. Score each criterion from 1 to 5 and weight by importance.
| Criterion | Weight | 1: Insufficient | 3: Acceptable | 5: Excellent |
|---|---|---|---|---|
| Technical transparency | 30% | Vague answers, marketing jargon | Clear but incomplete explanations | Detailed documentation, precise metrics |
| Error handling | 25% | “Our AI never makes mistakes” | Errors acknowledged without solutions | Documented limits and fallback mechanisms |
| Adaptability to customer data | 20% | Refuses your data | Testing possible, unclear terms | Structured POC on your data offered |
| Deployment maturity | 15% | No production customer | A few customers, unverifiable feedback | Verifiable references and case studies |
| Contract terms | 10% | Long commitment, unclear exit | Negotiable standard terms | Flexibility, exit provisions, clear SLAs |
Interpreting the Score
- 4.0–5.0: credible solution. Proceed confidently to POC.
- 3.0–3.9: potential, but unresolved issues. Request written clarification.
- 2.0–2.9: significant warning signs. Compare alternatives before committing.
- Below 2.0: move on. AI-washing risk is high.
The “We Will See in the POC” Trap
Do not substitute a POC for critical evaluation. Poor scoping—no predefined success metrics, unrepresentative data or insufficient duration—proves nothing. Define success before launch, not afterward.
A serious POC lasts at least 30 days, uses a representative sample of real rather than synthetic data, measures normal and edge cases, and compares with the existing process, with or without AI.
Build Your Own Test Scenario
Prepare Test Data Before the Demo
Do not let the vendor select demonstration data. Prepare three categories:
Easy data: standard cases any reasonably trained system should handle. This is the baseline; failure here makes further testing pointless.
Realistic data: representative daily cases with ambiguities, entry errors and heterogeneous formats. This is the real test.
Edge-case data: known difficult cases, including business exceptions, rare languages and unusual formats. The goal is not to trap the vendor but understand behavior at the system's limits.
Document Results Structurally
For every test, record input, expected output, actual output, response time and verdict: success, partial failure or complete failure. This factual table is worth more than subjective impressions.
Retain the documentation for vendor comparisons and as the POC reference.

Involve a Technical Specialist
You do not need to be a developer to ask these 12 questions. But when moving from POC to contractual commitment, an independent technical perspective is essential—not to endorse the choice, but to identify what the demo and salesperson left unsaid.
This might be a CTO, freelance solution architect or development partner able to audit architecture, APIs and documentation. A few person-days cost little compared with a failed AI project.
French Considerations
The AI Act and Its Evaluation Implications
The European AI Act, which entered into force in June 2024, classifies systems by risk. If your use case involves automated decisions affecting people in recruitment, credit or health, the vendor must demonstrate compliance with corresponding requirements.
Ask directly during the demo: “Which AI Act risk category applies to your system for our use case, and what compliance measures have you implemented?” Unfamiliarity or evasion is a major warning sign for any business operating in Europe.
The French Market in Numbers
French adoption context matters when evaluating local vendor maturity. According to INSEE 2024 data, 10% of French companies with more than 10 employees use at least one AI technology, up 4 points from 2023 but below the European average of 13%.
Adoption reaches 33% among companies with more than 250 employees and 42% in information and communication. The ecosystem contains over 1,000 AI startups and 16 AI-oriented unicorns. Mature solutions coexist with prototypes, making rigorous evaluation essential.
GDPR and Data Sovereignty
Every AI solution processing personal data must comply with GDPR. Beyond theoretical compliance, ask practical questions: where is data hosted, which subprocessors access data sent to the model, does it transit outside the EU, and can the vendor provide a compliant Data Processing Agreement (DPA)?
For sensitive sectors such as health, finance and defense, sovereignty extends beyond GDPR. If the model is hosted by a US hyperscaler, your data may be subject to the CLOUD Act. That is not always disqualifying, but must be documented and knowingly accepted.
After the Demo: Secure Your Decision
Require Written Commitments on Metrics
Put every verbal promise in writing. Request a document specifying committed accuracy, latency and availability, the conditions under which they were measured and contractual consequences if production performance falls short.
A vendor refusing written commitments on demonstrated performance sends a clear message: it cannot reproduce those results.
Structure the POC Around Go/No-Go Criteria
Before launch, define binary continuation criteria:
- Minimum accuracy on real data, for example > 90% on standard cases
- Maximum acceptable P95 latency, for example < 2 seconds
- Minimum availability during testing, for example > 99.5%
- Integration with at least one existing system, such as CRM or ERP
- Cost per transaction consistent with the business case
If any criterion is unmet at the end, the decision is no-go regardless of the original demo's quality.
Verify Customer References Independently
Vendor website case studies are marketing tools, not evidence. Ask to speak directly to production customers—not strategic partners or early adopters, but organizations using the product daily for at least six months.
Ask three questions: does the system deliver its promises, what problems arose, and would you recommend this vendor to a peer? Answers are worth more than any demo.
FAQ
How Can I Tell Whether an AI Demo Is Rigged? Ask to test your own data in real time. Refusal or delay is a warning. Watch whether the demonstration always follows the same path: a genuinely functional system can be tested ad hoc without preparation.
Do I Need Technical Skills to Evaluate an AI Demo? No. These 12 questions are designed for nontechnical people. Vendor answers are revealing: serious providers use figures and documentation rather than jargon or metaphors. Involve a technical specialist at POC stage.
How Long Should a Reliable AI POC Last? At least 30 days on real data is recommended. Shorter periods miss performance variation from data volume, edge cases and load increases. Define success metrics before launch, not after.
What Is AI Washing, and How Do I Protect Against It? It means claiming AI capabilities a product does not actually possess. Require documented metrics, inspect technical documentation and independently verify customers. The US SEC now penalizes this practice.
What Does the European AI Act Require of AI Vendors? The AI Act, which entered into force in June 2024, classifies systems by risk. For high-risk recruitment, credit and health systems, vendors must guarantee model explainability, technical documentation and human supervision. Ask which category applies to your use case.
What Budget Should I Allow to Evaluate an AI Vendor Properly? Vendors often offer POCs free or at reduced cost. The real cost is internal: preparing data, assigning evaluators and potentially engaging an independent technical specialist for a few days. Allow 5–15 person-days in total, a small investment against the risk of a six-figure contract for an unsuitable solution.
AI Coder Squad: Evaluating AI Also Means Knowing What You Want It to Do
Asking the right vendor questions requires a clear business need, defined performance criteria and anticipated integration constraints. That is exactly the work a senior development team does before building or integrating an AI component.
AI Coder Squad designs custom applications and AI agents for businesses that want to move quickly without sacrificing quality—with senior developers and an AI-powered approach.
→ Start your project and discover how AI Coder Squad can accelerate your next delivery.