Every AI vendor selling into public housing right now has a good demo. Ours included. Demos are built from clean data, on a happy path, by people who know exactly which button not to press.
That is not dishonest, it is just what a demo is. The problem is that almost every failure in this market is a data failure or an adoption failure, and a demo is structurally incapable of showing you either.
These are the questions we would ask. Some of them are uncomfortable for us to answer too, which is the point — a question only a vendor with real deployments can answer is a useful question.
On your data, not theirs
1. Can we run this on our own files, in a pilot, before we buy?
The single most informative question. Not a sandbox with sample households — a real caseload, your imports, your admin plan, your staff. If the answer involves a lengthy paid implementation before you see anything work on your data, you are buying on faith.
2. What happens when a record has no date of birth and no SSN?
This is the question that separates vendors who have processed real PHA imports from vendors who have not. Real PHA data is missing more than any product plan assumes. We designed our caller verification around DOB or SSN last-4 with fallback questions, then ran it against an authority's imported applicant data and found zero out of 112 applicants had either. Every call bottomed out in a quiz about the month someone applied eight years ago. We rebuilt it around the two things files reliably contain and residents reliably know: their name and the phone number on their file.
Ask the question. Watch whether the answer is specific.
3. How do you handle a household that appears in three places with three spellings?
Duplicate and near-duplicate records are the normal state of a PHA file. Ask how the system resolves identity, what it does when it is unsure, and — critically — what it does not do when it is unsure.
4. Show me the audit trail for one extracted number.
Take a paystub. Ask them to show you the extracted income figure, the source document with the line visible, the detected pay frequency, and the annualization method. If a specialist cannot verify the number faster than they could compute it themselves, the tool adds work.
On what it can actually do
5. Does the AI read live case state, or only policy documents?
A system grounded in a knowledge base of your policies can explain the recertification process. It cannot tell a caller that their packet is missing one recent paystub — and that is the call they actually made. Ask which one you are buying. Many products in this market are the first while being marketed as the second.
6. Can it complete an action mid-interaction, or only answer and route?
Answering, summarizing, and routing is a real product category and it has value. It is not the same as finishing. The call that ends "I've texted you a link, take a photo, you're done" is the one that prevents the next three calls. Ask specifically: can it send the link, attach the returned document to the correct case, and update the case state?
7. What does it do with a call that is not a resident?
Almost nobody asks this and everybody should. This summer an MFA robocall dialed one authority's main line and read a payroll-tax one-time passcode aloud to the AI agent, while the staff member who triggered it waited by a phone that never rang. The code sat in the transcript and nothing was watching for it. We now detect authentication codes spoken as digits, as words, and in NATO phonetics, and alert staff. Ask what their system does. The answer tells you how long they have been running real phone lines.
8. What happens on the calls it cannot handle?
Ask about the handoff, not the deflection rate. Does staff pick up with the transcript and context in front of them, or does the caller start over? Does a VAWA disclosure, a reasonable accommodation request, or a distressed caller reach a person quickly and reliably?
On the boundary
9. Which decisions does the software make, and which stay with our staff?
In HUD-regulated work there is one acceptable answer: eligibility, income determination, rent, HAP, and inspection outcomes are staff decisions. The software prepares; a person approves. Ask them to state it plainly, and ask whether it is enforced by the system's design or only by their policy. Those are very different guarantees.
10. Can a staff member override anything the AI produced, and is that logged?
Overrides should be trivially easy and completely logged. If overriding is awkward, staff will work around the system entirely, which is how tools die in month six.
11. What is your data handling — training, retention, PII?
Specifically: is our data used to train models (the answer should be no), what is encrypted in transit and at rest, how long are recordings and transcripts retained and who can delete them, and what PII does the voice agent ever speak aloud. Ask for their SOC 2 status and their alignment to NIST 800-53 in writing rather than as a claim on a slide.
On what it costs to actually adopt
12. Who does the configuration, and who sits with our staff in week two?
McKinsey's 2026 public-sector AI report suggests roughly $5 of adoption, training, and change management for every $1 of technology spend. Most proposals in this market are 90% license and a line that says training included. Ask for names and hours. Ask who rewrites the resident-facing language after the first 200 calls show that nobody understands the phrase "interim reexamination."
13. What does year two cost, and what drives the variance?
Per-minute pricing, per-case pricing, and per-seat pricing all behave differently when volume moves. Ask what happens to your bill during a HOTMA-driven notice surge, or in a month when call volume doubles. Ask what the exit looks like: can you export your data, in what format, and what does it cost?
14. Give me a reference at our size, with our program mix — and I am going to ask them what broke.
Not "are you happy with it." Ask their reference: what went wrong in the first 60 days, what did staff resist, what did you have to change about your own process, and what does the vendor still not do well. Every real deployment has answers to those questions. A reference with no complaints has not used the product much.
A note on procurement structure
The way to de-risk all of this is not a longer RFP. It is a smaller first commitment.
Scope the first engagement to one workflow and one measurable outcome — days from recert notice to signed 50058, or hours from RFTA received to blocker list issued — sized to fall under your small purchase rules and your admin plan's thresholds. Run it on real cases. Measure against a baseline you captured beforehand.
That structure is also, per McKinsey's data, the one that reaches production: programs scoped to a whole service journey get there about 70% of the time, versus about 30% for isolated use-case projects. A journey-scoped pilot with a real number at the end gives you either a defensible full procurement or a cheap, fast no.
We are happy to be evaluated this way. If you want to put us through these fourteen questions on a call, we will answer all of them, including the ones where the honest answer is "not yet."