For companies · AI readiness · Governance · Infrastructure

Most AI initiatives fail on the data, not the model.

Your model is fine. What breaks is undocumented tables, conflicting definitions, missing lineage, and permissions that stop applying three hops downstream. SchemaVita measures whether your data estate can carry what you're shipping.

The problem

The failure mode is upstream of the model.

Not a hunch — the most consistently reported finding in enterprise AI deployment, and measurable before you spend another quarter on a pilot.

60%
of AI projects will be abandoned by organizations that lack AI-ready data.
Gartner projection, through 2026
52%
of questions produced fabricated answers when the same RAG system ran on unvetted data. On curated content, near zero.
Published medical RAG study
+9.2
percentage-point gain in retrieval precision from metadata enrichment alone, with no architecture change.
Published RAG benchmark
Services

Two ways in, depending on where it hurts.

Both are real deliverables, priced on their own merits and never credited back against later work — that would make it a sales call rather than an audit.

For data leadership

AI Data Readiness Assessment

Your AI pilot has stalled and the data is the suspect. Seven dimensions measured against published standards — scorecard, findings register, and a dependency-ordered roadmap.

$5,500 — $18,000
1 to 5 weeks
CDO · Head of Data · VP Analytics
For engineering leadership

AI Development Governance Audit

Your team adopted Claude Code, Cursor, or Copilot under pressure and nobody knows what's entering the codebase. Standards drift, secret handling, review gates, and supply-chain exposure — measured, then a governance plan.

$7,500
2 weeks
CTO · VP Engineering · Head of Platform

AI Data Readiness Assessment

Fixed fee · 1–5 weeks

A measured verdict on whether your data can carry the AI system you intend to build. Scored on evidence rather than interviews, with every artifact handed over so your team can re-run the numbers.

  • Instrumented measurement across all seven dimensions
  • Scored against published standards — DAMA-DMBOK, NIST AI RMF, NIST SP 800-53
  • Findings tied to your use case — RAG, agentic, or predictive
  • Sequenced roadmap with owners, effort, and prerequisites
Focused
$5,500
1 week · ~12-page report
One data domain or a single AI use case. Two to three interviews.
Seed stage · small ops team
Standard
$9,500
2–3 weeks · ~30-page report
Full data estate, up to four source systems, five to eight interviews. The right scope for most stalled or pre-launch initiatives.
Series A–B · mid-market
Deep
$18,000+
4–5 weeks · ~60-page report
Multi-domain, regulated, or multi-entity estates. Compliance mapping included.
Regulated · PE portfolio

Implementation

Scoped from the roadmap · $18k–$120k

Building what the assessment says needs building. Fixed-fee against a defined scope, as a focused sprint or a full program.

  • Data governance for AI — documentation standards, quality contracts, ownership, lineage, and the monitoring that keeps them honest
  • Governed AI-assisted development — Claude Code, Cursor, or Copilot with enforced standards, hooks, security guardrails, and review gates
  • Agentic and self-healing pipelines — detect their own failures, quarantine bad data, recover without a human
  • Warehouse and transformation modeling — canonical entities, stable keys, AI-readiness built in

Ongoing

Retainer · from $2,000/month

For teams with the systems but not the specialist. Governance ownership, standards enforcement, and design review before new AI features ship.

  • Fractional AI infrastructure engineer — set days each month, from $5,000
  • Advisory retainer — async access and monthly review, $2,000
  • Quarterly readiness review — re-measure, report movement, flag regressions, $2,500 per quarter
Method

Seven dimensions, measured against published standards.

The framework derives from public sources, so you can check the work. Each dimension scores on a five-level maturity scale, weighted for AI readiness rather than general data health.

D1

Data Quality

Accuracy, completeness, consistency, timeliness, validity, uniqueness — profiled and measured, not asserted. Completeness and consistency carry extra weight.

DAMA-DMBOK six-dimension set
D2

Metadata & Documentation

Whether your estate is legible to a retrieval system. Accurate numbers aren't AI-ready if nothing can discover them or their provenance is unknown.

dbt project evaluator · dbt-coverage · catalog metrics
D3

Lineage & Traceability

Whether any value traces to its origin and forward to its consumers. Required for reproducibility, staleness detection, and GDPR, HIPAA, and EU AI Act obligations.

OpenLineage · catalog-native lineage
D4

Access, Entitlement & Sensitivity

Whether governance controls survive the trip downstream. Classification has to happen before indexing — repairing it at the vector store is too late.

NIST SP 800-53 · SOC 2 Trust Services Criteria
D5

Structure & Semantics

Whether the same concept looks the same everywhere. Canonical fields, stable identifiers, one agreed definition per metric.

Published AI-ready data literature
D6

Pipeline Reliability & Freshness

Whether the data arrives on time, and whether anyone finds out when it doesn't. Freshness, test pass rates, alerting coverage.

Source freshness and run artifacts
D7

AI Governance Posture

Cross-cutting. System inventory, ownership, risk categorization, evaluation practice, and incident response for model failure.

NIST AI RMF 1.0 — Govern, Map, Measure, Manage
How it runs

Measure first. Interview second.

Most assessments are interviews with a scorecard attached. This one instruments your estate first, so the conversations interpret evidence rather than collect opinions.

01

Scope

A short call to fix the AI use case. RAG, agentic, and predictive systems have different failure modes, and the weighting changes accordingly.

02

Measure

Scripted instrumentation against your warehouse and transformation layer — coverage, quality, lineage, entitlement, freshness.

03

Interpret

Interviews with owners, consumers, and whoever drives the initiative — to explain the numbers and score what artifacts can't show.

04

Deliver

Scorecard, findings, and a dependency-ordered roadmap, walked through with your decision-maker. No pitch.

Questions

Before you ask.

What access do you need, and how is it handled?

Read-only access to the warehouse and transformation layer, plus catalog access where one exists — no write permissions at any point. Covered by a mutual confidentiality agreement before any credential changes hands, and revoked at delivery.

Is the fee credited against implementation work?

No, and that's deliberate. Pricing the assessment on its own merits means I have no financial stake in what the findings say — including when the honest answer is that your data is in better shape than you feared.

How is this different from what a consultancy would sell us?

The measurement is instrumented rather than interview-led, so findings are reproducible — you get the artifacts and can re-run them in six months. And the framework derives from published standards, so you can check the reasoning rather than taking the method on trust.

What if the assessment says we are not ready?

Then it says so, and sets out what to do about it in sequence. The point of measuring is that the answer can go either way — and knowing early is far cheaper than another two quarters of pilot work.

Do you work with our existing team or replace them?

Work with them, always. The deliverable is written for your engineers to act on, and the roadmap assigns work by role so it can be picked up internally.

What are your availability and engagement limits?

I take a small number of engagements at a time rather than stacking them, and I'll tell you at the first call when I could realistically start.

About

Built by someone who does this at scale.

SchemaVita is Cordero Perez. I build the governance layer that makes AI systems accurate, reliable, and safe to deploy — documentation quality, entitlement monitoring, security guardrails, and the standards infrastructure that lets agentic systems act on organizational data without producing confident nonsense.

I do that work inside a Fortune 500 technology company, under real compliance pressure. Before that, three years as a senior consultant in AI and data engineering at a Big Four firm, and seven years in federal and municipal oversight.

The name means roughly structure brought to life — which is the job. Design the thing properly, then make it run.

Data & Analytics Developer
Fortune 500 technology · current
Senior Consultant, AI & Data Engineering
Big Four consultancy · 3 years
Data Analyst, Oversight & Investigations
Municipal government · 3 years
Certifications
AWS Cloud Practitioner · MIT/edX Supply Chain Analytics · MIT/edX Supply Chain Technology & Systems · Tableau Desktop Specialist · PCEP Python · SOA Exam P
Get in touch

Find out whether your data can carry it.

Tell me what you are trying to build and where it is stuck. If an assessment is not the right thing, I will say so — sometimes the answer is one conversation, not an engagement.

Contact us