Product
Platform AI agents Security & governance
Who it's for
Asset management Acquisitions Finance & IR
Company
About Resources PricingTrust Center Read the thesis Book a demo →
Field notes

How to evaluate AI for institutional real estate

Every vendor demo in this category looks the same: a document goes in, a fluent answer comes out, the room nods. The differences that will decide whether the tool survives contact with your auditors, your IC, and your rent roll are exactly the ones a demo cannot show. Here are the seven questions that surface them.

Built AI · July 2026 · ~9 min read
Built AI
Field notes · July 2026 · 9 min read
ProcurementDiligence
Key takeaways
  • Demos in this category are indistinguishable by design: extraction plus fluent prose is table stakes. The deciding differences are structural and invisible on a projector.
  • Seven questions expose the structure: provenance, determinism, document precedence, integration depth, tenancy and model policy, the approval workflow, and time-to-onboard.
  • The strongest evaluation is a proof of concept on your own portfolio, scored against a quarter you already know the answers to.
  • Whatever you choose, insist on zero autonomous external actions from day one. Trust is granted per workflow, not per vendor.

Procurement teams in institutional real estate are seeing more AI pitches in a quarter than they used to see in a career, and the pitches have converged. Each one drops a lease or an offering memorandum into a chat window, and each one produces a summary, a table, and a confident answer to a question. On a projector, a wrapper around a general-purpose model and a purpose-built platform look identical. The budget, unfortunately, cannot tell the difference either, which is how firms end up two renewals deep into a tool their auditors will not accept.

Why demos all look the same

Extraction and fluent prose are commodities now. Any competent team can wire a frontier model to a document store and demonstrate an impressive twenty minutes. What a demo cannot show is what happens on the thousandth document, on the amendment that contradicts the original, on the figure that gets challenged in committee, or on the security questionnaire your LP sends. Those are structural properties, and structure has to be interrogated, because it will never volunteer itself in a demo.

A demo shows you the system on its best day. Diligence is about its worst day: the contradicted lease, the challenged figure, the audit three years out.

The seven questions

1. Where did this number come from? Show me without leaving the screen.

Click any figure in the output and demand the source: the clause, the GL line, the assumption. Then change an input and confirm the figure recomputes and the trail updates. If the system explains instead of tracing, the citations are generated, not structural, and they will not survive an audit.

2. If I run it twice, do I get the same answer?

Financial figures must be deterministic: computed on an engine, not sampled from a model. Ask the vendor to run the same package twice and diff the numbers. Prose can vary; the DSCR cannot. A system that cannot promise this cannot produce a covenant certificate.

3. What happens when documents disagree?

Real portfolios are hierarchies: original lease, three amendments, a side letter, an estoppel. Ask how the system decides which term governs, and whether it can show you both the superseded and the governing value. Flat extraction that treats every document as equally true fails precisely on the assets where the money is.

4. Does it read my systems, or do I feed it?

If analysts have to export, upload, and re-upload, you have bought another manual step. The system should read Argus, Yardi or MRI, and your Excel models where they live, keep them reconciled continuously, and treat your systems of record as exactly that. Writes, if any, happen only on human approval.

5. Where does our data live, and what trains on it?

The acceptable answers are specific: your tenant, optionally your cloud or on-premise, no training on your data, bring-your-own-model if your policy requires it. Vague answers about "enterprise-grade security" without a deployment diagram are a no.

6. What can it do without a human?

The right answer for regulated capital is: nothing external. Drafts, computations, and alerts happen autonomously; sends, filings, and writes wait for a named person. Ask to see the audit log of an agent run, including the blocked actions.

7. How long until our book is live?

Rip-and-replace platforms measure onboarding in quarters and migrations. A layer that sits on top of your stack should measure it in weeks, with a real reference: portfolios of hundreds of assets onboarded without a migration. Ask for the customer who did it.

Scorecard
Seven questions, and what a pass looks like
Provenance
Every figure traces to a clause, line, or assumption, live on screen. Pass: click-through, not explanation
Determinism
Same inputs, same output, every run. Pass: a diff of two runs is empty
Precedence
Amendments and side letters supersede correctly. Pass: shows governing and superseded terms
Integration
Reads Argus, Yardi, MRI, and Excel in place. Pass: no export-upload loop
Tenancy
Your tenant, self-host option, no training on your data. Pass: a deployment diagram
Approval
Zero autonomous external actions, logged. Pass: the audit log shows the blocks
Onboarding
Weeks, no migration, referenceable. Pass: a named customer at scale
Print this and take it into the second meeting. A vendor with real infrastructure will enjoy these questions; a wrapper will redirect to the roadmap.

Red flags and pass signals

Some answers end the conversation politely. "The model explains its reasoning" in response to the provenance question means citations are generated. "Results may vary slightly between runs" means no deterministic engine, which means no auditor. "We are working on Yardi" means your analysts are the integration. And any hesitation on the training question means the data governance review will fail anyway, so save everyone the quarter.

The one-sentence test

Ask the vendor to complete this sentence: "When your auditor asks how the machine produced this number, you will show them ______." If the answer is a replayable calculation over cited sources, keep talking. If it is a confidence score, stop.

Run the proof on your own book

The final filter is empirical, and it is the one that removes the guesswork: a proof of concept on a slice of your real portfolio, in your environment, scored against a period you already know the answers to. Pick a quarter that had a known variance, a covenant that tightened, a deal you underwrote by hand. Let the system reproduce the work, then compare: the numbers, the trails, and the hours. The tools that survive this are infrastructure. The rest were demos.

  • Score outcomes, not features. Time to a defensible memo, breaks caught in reconciliation, days of warning on the covenant.
  • Involve the skeptics. The controller and the analyst who will live in it, not just the innovation team.
  • Keep the bar institutional. Output your IC would sign, on data your auditor could replay.

These seven questions are, not coincidentally, a description of how Built AI is built: a knowledge graph with document precedence, a deterministic engine, citation by construction, native reads of Argus, Yardi, MRI, and Excel, multi- or single-tenant deployment with no training on your data, zero autonomous external actions, and onboarding measured in weeks. We would genuinely rather you asked every vendor all seven, including us, on your own book. See how our proof of concept works or book a walkthrough.