For companies · AI readiness · Governance · Infrastructure

Most AI initiatives fail on the data, not the model.

Your model is fine. Your retrieval architecture is fine. What breaks is undocumented tables, conflicting definitions, missing lineage, and permissions that quietly stop applying three hops downstream. SchemaVita measures whether your data estate — and the way your engineers build with AI — can actually carry what you are shipping.

The problem

The failure mode is upstream of the model.

This is not a hunch. It is the most consistently reported finding in enterprise AI deployment, and it is measurable before you spend another quarter on a pilot.

60%
of AI projects will be abandoned by organisations that lack AI-ready data.
Gartner projection, through 2026
52%
of questions produced fabricated answers when the same RAG system ran on unvetted data. On curated content, hallucinations fell to near zero.
Published medical RAG study
+9.2
percentage-point gain in retrieval precision from metadata enrichment alone, with no change to the retrieval architecture.
Published RAG benchmark
Services

Two ways in, depending on where it hurts.

Both are real deliverables, priced on their own merits, and neither is credited back against later work — that would make it a sales call rather than an audit.

For data leadership

AI Data Readiness Assessment

Your AI pilot has stalled and the data is the suspect. Seven dimensions measured against published standards, delivered as a scorecard, a findings register, and a dependency-ordered remediation roadmap.

$5,500 — $18,000
1 to 5 weeks
CDO · Head of Data · VP Analytics
For engineering leadership

AI Development Governance Audit

Your team adopted Claude Code, Cursor, or Copilot under pressure and nobody knows what is entering the codebase. Standards drift, secret handling, review-gate coverage, and supply-chain exposure — measured, then a governance plan.

$7,500
2 weeks
CTO · VP Engineering · Head of Platform

AI Data Readiness Assessment

Fixed fee · 1–5 weeks

A measured verdict on whether your data can carry the AI system you intend to build. Scored on evidence rather than interviews, with all measurement artifacts handed over so your team can re-run the numbers themselves.

  • Instrumented measurement of quality, coverage, lineage, entitlement, structure, and freshness
  • Scored against published standards — DAMA-DMBOK, NIST AI RMF, NIST SP 800-53
  • Findings tied to your specific use case, whether RAG, agentic, or predictive
  • Sequenced remediation roadmap with owners, effort, and prerequisites
Focused
$5,500
1 week · ~12-page report
One data domain or a single AI use case. Two to three interviews.
Seed stage · small ops team
Standard
$9,500
2–3 weeks · ~30-page report
Full data estate, up to four source systems, five to eight interviews. The right scope for most companies with a stalled or pre-launch initiative.
Series A–B · mid-market
Deep
$18,000+
4–5 weeks · ~60-page report
Multi-domain, regulated, or multi-entity estates. Compliance mapping included.
Regulated · PE portfolio

Implementation

Scoped from the roadmap · $18k–$120k

Building what the assessment says needs building. Fixed-fee against a defined scope, as a focused sprint or a full programme.

  • Data governance for AI — documentation standards, quality contracts, ownership, lineage, and the monitoring that keeps them true
  • Governed AI-assisted development — Claude Code, Cursor, or Copilot with enforced standards, hooks, security guardrails, and review gates
  • Agentic and self-healing pipelines — pipelines that detect their own failures, quarantine bad data, and recover without a human
  • Warehouse and transformation modelling — canonical entities, stable keys, semantic consistency, AI-readiness built in

Ongoing

Retainer · from $2,000/month

For teams that have the systems but not the specialist. Governance ownership, standards enforcement, reliability, and design review before new AI features ship.

  • Fractional AI infrastructure engineer — set days each month, from $5,000
  • Advisory retainer — async access and monthly review, $2,000
  • Quarterly readiness review — re-measure, report movement, flag regressions, $2,500 per quarter
Method

Seven dimensions, measured against published standards.

The framework derives from public sources rather than invention, so you can check the work. Each dimension scores on a five-level maturity scale, weighted for AI readiness specifically rather than general data health.

D1

Data Quality

Accuracy, completeness, consistency, timeliness, validity, uniqueness — profiled and measured, not asserted. Completeness and consistency carry extra weight, because missing or conflicting data is what produces unreliable models.

DAMA-DMBOK six-dimension set
D2

Metadata & Documentation

Whether your estate is legible to a retrieval system. A table of accurate numbers is not AI-ready if nothing can discover it, its columns are cryptic, or its provenance is unknown.

dbt project evaluator · dbt-coverage · catalog metrics
D3

Lineage & Traceability

Whether any value traces to its origin and forward to its consumers. Required for reproducing model results, for staleness detection, and for GDPR, HIPAA, and EU AI Act obligations.

OpenLineage · catalog-native lineage
D4

Access, Entitlement & Sensitivity

Whether governance controls survive the trip downstream. Sensitivity classification has to happen before indexing — repairing access control at the vector store is already too late.

NIST SP 800-53 · SOC 2 Trust Services Criteria
D5

Structure & Semantics

Whether the same concept looks the same everywhere. Canonical fields, stable identifiers, normalised units and time zones, one agreed definition per metric across teams.

Published AI-ready data literature
D6

Pipeline Reliability & Freshness

Whether the data arrives, on time, and whether anyone finds out when it does not. Freshness, test pass rates, alerting coverage, and how often a human intervenes by hand.

Source freshness and run artifacts
D7

AI Governance Posture

Cross-cutting. System inventory, ownership, risk categorisation, evaluation practice, incident response for model failure, and third-party model visibility.

NIST AI RMF 1.0 — Govern, Map, Measure, Manage
How it runs

Measure first. Interview second.

Most assessments are interviews with a scorecard attached. This one instruments your estate before anyone gets asked a question, so the conversations interpret evidence rather than collect opinions.

01

Scope

A short call to fix the AI use case in question. RAG, agentic, and predictive systems have different failure modes, and the weighting changes accordingly.

02

Measure

Scripted instrumentation against your warehouse and transformation layer. Coverage, quality profiling, lineage completeness, entitlement inheritance, freshness.

03

Interpret

Interviews with owners, consumers, and whoever drives the initiative — to explain the numbers and score what artifacts cannot show.

04

Deliver

Scorecard, findings, and a dependency-ordered roadmap, walked through with your decision-maker. No pitch in that meeting.

Questions

Before you ask.

What access do you need, and how is it handled?

Read-only access to the warehouse and transformation layer, plus catalog access where one exists. No write permissions at any point.

Everything is covered by a mutual confidentiality agreement before any credential changes hands, and access is revoked at delivery.

Is the fee credited against implementation work?

No, and that is deliberate. Pricing the assessment on its own merits means I have no financial stake in what the findings say.

If the honest answer is that your data is in better shape than you feared, I want to be able to tell you that plainly.

How is this different from what a consultancy would sell us?

Two things. The measurement is instrumented rather than interview-led, so the findings are reproducible — you get the artifacts and can re-run them yourselves in six months.

And the framework derives from published standards, so you can check the reasoning rather than taking the method on trust.

What if the assessment says we are not ready?

Then it says so, and it sets out what to do about it in sequence. The point of measuring is that the answer can go either way.

Knowing early is considerably cheaper than discovering it after another two quarters of pilot work.

Do you work with our existing team or replace them?

Work with them, always. The deliverable is written for your engineers to act on, and the remediation roadmap assigns work by role so it can be picked up internally.

What are your availability and engagement limits?

I take a small number of engagements concurrently rather than stacking them, and I will tell you honestly at the first call when I could realistically start.

About

Built by someone who does this at scale.

SchemaVita is Cordero Perez. I build the governance layer that makes AI systems accurate, reliable, and safe to deploy — documentation quality systems, entitlement monitoring, automated security guardrails feeding into AI build workflows, and the standards infrastructure that lets agentic systems act on organisational data without producing confident nonsense.

I do that work today inside a Fortune 500 technology company's finance organisation, under real compliance pressure. Before that, three years as a senior consultant in AI and data engineering at a Big Four firm, and three years as a data analyst for a municipal oversight and investigations division, where the work was cited in the New York Times.

The name means roughly structure brought to life — which is the job. Design the thing properly, then make it run.

Data & Analytics Developer
Fortune 500 technology · Finance · current
Senior Consultant, AI & Data Engineering
Big Four consultancy · 3 years
Data Analyst, Oversight & Investigations
Municipal government · 3 years
Certifications
AWS Cloud Practitioner · MIT/edX Supply Chain Analytics · MIT/edX Supply Chain Technology & Systems · Tableau Desktop Specialist · PCEP Python · SOA Exam P
Get in touch

Find out whether your data can carry it.

Tell me what you are trying to build and where it is stuck. If an assessment is not the right thing, I will say so — sometimes the answer is one conversation, not an engagement.