This website is constantly evolving as we learn how to explain evals more clearly.

Early test run · 27 August 2026

How 12 AI models handled the same resident conversations.

Each model answered 30 real resident situations, one reply each. Bars show the share that passed.

Share of replies that passed

Passed replies meet every required behavior and at least 3 of 4 writing checks. Errors such as invented facts or promises fail the reply.

100
75
50
25
OpenAI
GPT-5.6 Sol9.68¢
OpenAI
GPT-5.6 Luna1.07¢
OpenAI
GPT-5.6 Terra6.99¢
Z.ai
GLM 5.24.84¢
Meta
Meta Muse Spark 1.220.23¢
Anthropic
Claude Opus 530.40¢
Anthropic
Claude Sonnet 59.47¢
Moonshot AI
Kimi K310.63¢
Google
Gemini 3.7 Flash8.24¢
Xiaomi
MiMo V2.50.92¢
DeepSeek
DeepSeek V4 Flash 07310.85¢
NVIDIA
Nemotron 3 Ultra 550B A55B3.81¢

Cost for all 30 repliesUS cents · generation only

Cost details

Provider-reported costs to generate the replies on 27 August 2026. AI grading and calibration are excluded. These are inference costs, not the separate OpenRouter charges, which were zero for some models. They are not current price quotes.

Click a bar for a score. Replies are below.

Next: public tests of multifamily AI.

Next: public tests of multifamily AI.

We’ll test resident engagement software and publish the tests, scoring rules, and results for anyone to check.

We’ll test resident engagement software and publish the tests, scoring rules, and results for anyone to check.

PLANNED WORKFLOW

Tests haven’t run yet

Tests haven’t run yet

AppFolio AI

AppFolio AI

Not tested yet

Not tested yet

BetterBot

BetterBot

Not tested yet

Not tested yet

EliseAI

EliseAI

Not tested yet

Not tested yet

ResiDesk

ResiDesk

Not tested yet

Not tested yet

Yardi AI

Yardi AI

Not tested yet

Not tested yet

Public tests

Same scoring rules

Public tests

Same scoring rules

Public results

Anyone can inspect

Public results

Anyone can inspect

Public results

Anyone can inspect

Planned by ResiDesk. Listing a vendor does not imply participation or endorsement.

Why

Why evals matter.

AI speaks for your team. A missed detail can change the meaning.

Resident’s request

“Only enter while I’m home.

What the eval checks

Does the reply keep that limit?

The model can change behind the same inbox.

Resident inboxSame interface
Resident
Thursday or Friday works. Please only enter while I’m home.
Model A
We’ll confirm a time, with entry only while you’re home.
Model B
We’ll confirm a time.

Entry limit missing

Illustrative replies. Not test results.

Test the replies again.Check that the entry restriction survives the switch.

Instructions and tools can change too.

Across thousands of conversations, small omissions are hard to spot.

What

What is an eval?

An eval tests model replies using real conversations.

Excerpts · Translated · Messages omitted

Resident
This is how my kitchen has been for 1 year.
Associate
Also, please confirm if we have your permission to enter for possible repairs or assessment even when you're not home.
Resident
Only when I am at home will I be at home on Thursday, Friday.
Example replies

No visit promised

Thursday or Friday may work for you.

Unbooked visit promised

maintenance will come Thursday or Friday.

Illustrative reply excerpts, not saved model replies. Nothing sent.

How

How we build an eval.

Use a real conversation to define what a good reply must do.

Resident’s reply

Only when I am at home (checklist item 2) will be at home on Thursday, Friday. (checklist item 3)

Checklist for this case

  1. Note with the work-order request

  2. Enter only while home

  3. Keep days as availability

  4. Do not re-ask for accessThe resident already answered.

Translated excerpt. These details become checks; the full conversation supplies the rest.

Grade every model with the same checklist.

What makes a reply pass?

  1. It respects limits and invents nothing.

    No invented facts, actions, or appointments. No ignored entry limits.

  2. It includes every required behavior.

    All four items in the checklist, in any wording.

  3. It meets at least three writing checks.

    Short, warm, no internal jargon, sounds like a text message.

All three are required. Draft criteria.

Grading details

List required behaviors first. The accepted human reply is a reference, not the only right wording.

Accepted human reply · Hidden from the AI during the test
Got it! I'll leave a note in the work order that maintenance can only enter when you're home on Thursday or Friday. Let us know if you need help with anything else.

Entry conditions from the conversation: maintenance may enter only while the resident is home. Thursday and Friday are availability, not a booked visit.

Hard failures fail a reply outright: ignoring entry limits or inventing facts, actions, or appointments.

Writing checks: short, warm, no internal jargon, sounds like a text message.

Draft threshold: 0 hard failures, 4 of 4 required behaviors, at least 3 of 4 writing checks. One missed writing check is allowed; one missed required behavior is not.

The resident later said thanks. That supports the human reply, not proof AI booked a visit or fixed the floor.

Stage 2: Cover the resident journey.

Build cases for every resident touchpoint.

During residence
  1. Prospect
  2. Leasing
  3. Billing
  4. Maintenance
  5. Renewals

30 cases. Not full coverage.

Stage 3: Make results public.

Let others inspect the cases, replies, and grades.

  1. CaseExcerpts
  2. ChecklistExample
  3. Replies48 / 360
  4. Grades48 / 360

Partial draft. Public does not mean independently checked.

Research details
Conceptual: what each test type could publish. This draft is the second row.
Type of testCaseRulesAll repliesGradesWho checks
A demoShown as publishedShown as not publishedShown as not publishedShown as not publishedAI builder
This draftExcerptsExample48 / 36048 / 360AI builder
A complete public evalShown as publishedShown as publishedShown as publishedShown as publishedAI builder
With an outside evaluatorShown as publishedShown as publishedShown as publishedShown as publishedOutside evaluator

ResiDesk builds AI and ran this research; not independent certification.

Stage 4: Repeat the test.

Check consistency. Retest after model, instruction, or tool changes.

  1. 27 August 2026Development run
  2. Not yet runSame cases, unchanged
  3. Not yet runAfter a change

Create a free website with Framer, the website builder loved by startups, designers and agencies.