# After-Hours Answer Study: protocol deviation log (public copy)

Published with the results at https://www.elevyr.com/research/after-hours-results. This copy is the working log word for word, except that community numbers, names, license numbers, phone numbers, call ids, account emails and the caller's name are withheld so that no community can be traced to its own result (protocol 12.2, DEV-017), and one personal detail is reworded.


**Published in full with the study.** Governed by `01-protocol.md` 11.2.

Real protocols have deviations. The dishonest ones are the deviations that go unlogged. A study with
an empty deviation log and 300 calls is less believable than one with four honest entries.

## Rules

1. **Every departure from the protocol gets a row.** However small. However embarrassing.
2. **Log it when it happens**, not at write-up. A deviation remembered three weeks later is a
   deviation reconstructed.
3. **Reference the `DEV-nnn` id in the affected call row** (`protocol_deviation` column).
4. **Never fix a deviation by editing the protocol.** The protocol is registered. It does not move.
5. **A definition that maps an outcome to no code, or to two codes, is not a deviation. It halts
   fielding.** See `01-protocol.md` 11.3.

## Log

| ID | Date | Section | What happened | Why | What was done | Calls affected |
|---|---|---|---|---|---|---|
| **DEV-001** | 2026-07-31 | 6.1, 7.4 | **The caller changed from a blind VA contractor to Ed Brancheau, the study's designer.** Post-registration, pre-fielding (zero calls placed). | Two contracted VAs went unresponsive, costing two field weeks (07-19 and 07-26 both lapsed with zero calls placed). Ed ruled 2026-07-31 that indefinite slippage costs the study more than a disclosed limitation does. **Note the registered protocol is internally inconsistent on this exact point:** §5.4 already states "Ed is the rater and cannot unsee individual calls, so a blinding claim would be a lie," while §6.1 and §7.4 specify a blind contractor. Fielding resolves that contradiction toward §5.4. | **The §7.4 blind-caller mitigation is gone, and the publication must say so plainly and near the top: the study's designer, who is also the vendor selling a product these numbers would justify, placed and coded every call.** The defense falls back entirely to §7.1 (the primary outcome is binary and near-observer-independent), the pre-registered definitions frozen before call #1, and the published row-level data. **Partially offset:** the "one offshore voice" limitation in §7.4 no longer applies, since the caller is now a local US voice closer to the family a real inquiry represents. The OSF registration is frozen and cannot be edited (rule 4 above), so the discrepancy is disclosed here rather than patched. **The public method page must be updated to match before call #1**, since it currently promises a blind contractor. | 0 |
| **DEV-002** | 2026-07-31 | 5.2 | Field week pushed a second time: field-start Sunday 2026-07-26 → **2026-08-02** (C Sun 08-02, B Tue 08-04, A Wed 08-05). | The 07-26 week lapsed with zero calls for the same reason 07-19 did: no caller was available. | `00-draw.js` re-run with the same registered seed (`0203102737`) and the same roster snapshot. **Verified: the `community_id` → cell → `dial_order` mapping is identical across all 300 calls before and after; only the three date fields moved.** `covariates.csv` untouched (freeze guard held). August 2026 contains no US federal holiday, satisfying the no-holiday-week rule in §5.2. | 0 |
| **DEV-003** | 2026-07-31 | runbook §1 | Calling line resolved: **a free Google Voice number with a California area code, call recording OFF**, account controlled by Ed. Closes the open loop left by the 07-19 Quo cancellation. | The 07-19 entry promised a replacement line would be "logged here once picked," and it never was. The Quo trial was cancelled 07-19 before it billed, releasing [number withheld]. | Google Voice is free, provides ONE consistent US caller ID, records no audio (California two-party consent), and Ed controls the login so callbacks reach him and the number can be abandoned after fielding. This satisfies the design requirement the registered protocol's "a single Quo line" phrase encodes: one consistent US number with recording off. | 0 |
| **DEV-004** | 2026-08-02 | 5.2 | Field week pushed a third time: field-start Sunday 2026-08-02 → **2026-08-09** (C Sun 08-09, B Tue 08-11, A Wed 08-12). Zero calls had been placed under the 08-02 date. | The caller (DEV-001) could not field the session the evening of 2026-08-02. Fielding 100 live calls without full attention risked coding quality on the one part of the study that depends on rater care. | `00-draw.js` re-run with the same registered seed (`0203102737`) and the same roster snapshot, `--field-start 2026-08-09`. **Verified: the `community_id` → cell → `dial_order` mapping in `schedule.csv` is byte-identical across all 300 rows before and after (columns 1-5); only the two date columns moved.** `covariates.csv` output from the re-run is byte-identical to the frozen file (diff empty), so the hand-coded `publishes_rate`/`rate_band`/`prospect_overlap` values already recorded for 39 communities are untouched. August 2026 contains no US federal holiday, satisfying the no-holiday-week rule in §5.2. | 0 |
| **DEV-005** | 2026-09-20 | 6.1, 7.4 | **The caller changed again, from Ed Brancheau to a paid third-party caller (a university student, name withheld in this public copy).** Zero calls had been placed under Ed as caller (DEV-001). | DEV-001 made Ed the caller because two contracted VAs went unresponsive. A hiring funnel built 2026-09-13 (`hiring-funnel.md`) opened a paid research-assistant role, and a university student applied and cleared all five screening gates on 2026-09-14. Ed hired her the same evening. | **This repairs most of the DEV-001 disclosure rather than adding a new limitation.** DEV-001's heavy disclosure read: "the study's designer, who is also the vendor selling a product these numbers would justify, placed and coded every call." The new caller is neither the designer nor the vendor, so that specific credibility problem is gone. Full sponsor-blinding still is not achieved (she is paid by and knows she was hired by Ed), but hypothesis-blinding is intact: the job posting and instructions never named Sloane, AI, or the study's expected finding. **The public method page must be updated to match before call #1**, since it currently still describes Ed placing the calls himself. | 0 |
| **DEV-006** | 2026-09-23 | 5.2 | Field week pushed a fourth time: field-start Sunday 2026-08-09 → **2026-09-27** (C Sun 09-27 20:00 to 22:05, B Tue 09-29 20:00 to 22:05, A Wed 09-30 10:00 to 12:05). Zero calls had been placed under the 08-09 date. | The 08-09 week lapsed with no call placed, and the caller then changed (DEV-005). A hired caller needed the new dates confirmed against her class timetable, which was done 2026-09-20. | `00-draw.js` re-run 2026-09-23 with the same registered seed (`0203102737`), the same roster snapshot (`roster-cdss-rcfe-2025-05-25.csv`, `--retrieved-on 2026-07-17 --portal-modified 2025-05-27`) and `--field-start 2026-09-27`. **`covariates.csv` is byte-identical to the frozen file (SHA256 `3b9f31816343f9627dd7311ec5841f9857de4a73fd7a3356d3b68afd1b2de49f` before and after)**, so the sample and the hand-coded values are untouched; `schedule.csv` carries 300 rows, 100 per session date. Pre-regen files kept in `_backup-pre-regen-20260923/`. The week of 2026-09-27 contains no US federal holiday, satisfying the no-holiday-week rule in §5.2. | 0 |
| **DEV-007** | 2026-09-24 | runbook §1 | Calling line changed from the DEV-003 Google Voice number to **Quo, Starter plan (paid, one user), number [number withheld]** (inbox name "Researcher"), account [account withheld] controlled by Ed. The caller dials through that same login. Zero calls had been placed on any line. | Google Voice hit an account limit (recorded 2026-09-15) and was never usable. The earlier Quo account ([account withheld], [number withheld]) was cancelled 2026-07-19. **A first number on the new account, [number withheld], was dropped the same day before any use:** Google's caller spam filter labeled it spam on Ed's own phone. Quo does not allow a number swap during a free trial, so the account was moved to the paid Starter plan and the number replaced. The replacement [number withheld] had zero reports on Robokiller and no web complaints, but **Google's filter labeled it spam too**, which shows the label attaches to new internet (VoIP) numbers in general rather than to either number's history. | ONE consistent US caller ID with a San Diego area code. Auto-record calls and transcription were switched OFF during setup (the free trial had turned recording ON by default); the paid Starter plan does not include auto-recording. AI assistant (Sona) not used; missed calls go to plain Quo voicemail; no call forwarding. The caller must not use Quo's manual record button. Registration with the Free Caller Registry (the database behind the AT&T, T-Mobile and Verizon spam labels) was attempted 2026-09-24 and failed on the site's own error; it was not retried, because it would take effect days into a six-day line and would not affect the Google label actually observed. The number is therefore unregistered. **Disclosed limitation:** some recipients, mainly on Google phone services, may have seen a spam label on the study's caller ID, which could lower pickup; a real family calling from a personal cell would not carry that label. The subscription is to be cancelled after the last session Wed 2026-09-30, releasing the number. This again satisfies the design requirement behind the registered "a single Quo line" phrase: one consistent US number with recording off. | 0 |
| **DEV-008** | 2026-09-24 | 4.2, codebook decision tree | **Coding rule added for call-screening assistants**, before any call was placed. Phones increasingly answer unknown callers with an automated screener ("say your name and I'll see if this person is available"). The registered codebook had no code for this, and 11.3 halts fielding on an outcome with no code. | Found 2026-09-24 when a test call from the study line to a personal phone reached exactly such a screener. | Rule: a screener is NOT an answer. The caller says "A family member looking into assisted living" (true, no invented name, per the script), waits, and codes whatever comes next with the existing tree (human, voicemail, and so on), noted "screened". Nothing within 30 seconds of speaking codes ringout / NO, noted "screened". Added to `02-codebook.md` and `05b-va-caller-sheet.md`. **Why this is fair to the communities:** a real family calling after hours is also an unknown number, so a screener is part of the real condition the study measures. The "screened" note lets a reader recompute every result with screened calls excluded. | 0 |
| **DEV-009** | 2026-09-24 | 3, 5.2, 04-analysis | **Two pairs of sampled communities share one phone line**, found while coding covariates, before any call. In one pair, both licenses belong to one building: same street address, same capacity, same phone, most likely an old and a new license both listed LICENSED in the 2025-05-25 roster. The other pair is two differently named communities that list the same phone number. Each pair is dialed twice in every session. | The frame is built from license records, and one building or one front desk can hold two licenses. The draw could not see that. | **The sample is NOT changed** (rule 4: the registration is frozen). Every slot is dialed as scheduled. Disclosed limitation: the second call to a shared line in the same session is not independent of the first. Pre-committed handling: the primary rate is reported as registered, plus a sensitivity rate with the later-dialed member of each shared-line pair dropped per session. | 0 |
| **DEV-010** | 2026-09-24 | 3.2, 11.6, 12.1, codebook covariates | **Covariates coded and frozen before call #1, with five reading rules made explicit.** Coded 2026-09-24 by AI research agents working only from each community's own website, blind to outcomes (no call had been placed). Result: publishes_rate YES 34, NO 48, UNCLEAR 18; rate_band <6000 25, >=6000 9; prospect_overlap TRUE 45. A spot check found a missed price on an inner page, so every NO and UNCLEAR row (64) got a second, deeper pass (inner pages, linked PDFs, raw page code); it changed 13 rows. Frozen file SHA256 388dcc7f4e80999e1042dadf8eb63ee5e8e59a8acdfb4337fd109192ef144fcc (was 3b9f3181... with the four columns blank); every other column verified unchanged. | The registered codebook left five cases open. | Rules applied, fixed before any call: (1) a price visible only after submitting a form or chat is not published (UNCLEAR); (2) "contact us" or "call for rates" is NO, while a pricing section or "starting at" wording with no figure is UNCLEAR; (3) an own site that could not be read at all is UNCLEAR (two communities); (4) a published daily rate is converted at 30 days (one community); (5) the lowest starting rate anywhere on the own site is used, per the codebook, even when it is independent living or memory care rather than assisted living (four communities). **prospect_overlap source:** the people ledger has no company column, so company names were taken from `sourcing-reservoir.csv` (the ledger's identity source), snapshot `ledger-snapshot-2026-09-24.csv`, 613 distinct names. The pipeline proposed 56 matches; each was reviewed at operator level and 11 rejected as word-only matches. Working files: `research-worklist.csv`, `prospect-overlap-review.csv`. | 0 |
| **DEV-011** | 2026-09-25 | 4.1, 11.3, 04-analysis 1.3, codebook decision tree | **Coding rule, a post-field check, and two sensitivity analyses added for carrier call blocking and spam labels**, before any call was placed. Adds rules only: the sample, the three primary outcomes, and the `outcome_subtype` list are unchanged. | DEV-007 records that the study line [number withheld] is an unregistered VoIP number that Google's filter labels spam. Two risks follow. (1) **The denominator risk, the larger one.** Carriers can refuse a labeled number with a recording ("not accepting calls", "your call has been declined") that a caller could easily code `dead_number`, the only exclusion. Each such miscode silently removes a real NO from the denominator, and from the paired McNemar test (04-analysis 1.3 drops any pair with an EXCLUDED call), which would flatter after-hours coverage. The registered codebook had no rule separating a block from a dead line. (2) **The label risk.** A label may lower pickup. It is held constant across cells (one number, the same 100 communities in every session), so the paired contrast absorbs it; but a rate read on its own does not. **Registering or swapping the number mid-week was rejected:** a label that clears partway through would lift pickup in the last session only (A, the weekday control), widening the day-versus-night gap in the direction the sponsor benefits from, and the timing of the change could not be observed. The line stays as-is through Wed 2026-09-30. | **(1) Coding rule** (added to `02-codebook.md` and `05b-va-caller-sheet.md`): a recording saying the call was blocked, declined, rejected or "not accepting calls" is `ringout` / NO, noted "blocked". `dead_number` only for a recording saying not in service, disconnected, or no longer in service. Any recording the caller cannot place: `ringout` / NO, noted "blocked". **(2) Post-field check:** after the last session ends Wed 2026-09-30 (so no check call can touch a live session), every number with a `dead_number` or "blocked" row is dialed once from a second phone not on the study line, hanging up at the first ring. Out of service for the second phone too: stays or becomes `dead_number`. Rings for the second phone: stays or becomes `ringout` / NO, noted "blocked". Any change is an appended correction row (`corrects_call_id`, this DEV id), never an edit. Results of every check call are logged below this row. **(3) Pre-committed sensitivity analyses,** labeled secondary, reported alongside and never instead of the registered primary: (a) every cell rate and the paired test recomputed with calls noted "blocked" or "screened" dropped; (b) the paired between-cell differences named in the write-up as the check that caller-ID labeling cannot drive, because the label is identical in every cell. Descriptive only, labeled exploratory: calls resolving to voicemail at 0 or 1 rings (`rings_to_resolution`), the pattern of a phone set to silence unknown callers. **Direction of any residual bias:** sessions run C (Sun), B (Tue), A (Wed, control). If recipients block the number after the first call, the control is hit hardest, which shrinks the day-versus-night gap rather than inflating it. | 0 |
| DEV-011 check results | 2026-09-30 | DEV-011 (2) | **Post-field check, run after the last session closed.** Five communities had at least one `dead_number` call; no call was noted "blocked" or "screened". Each number was dialed once on Wed 2026-09-30 about 15:05 PT from Ed Brancheau's personal cell, not the study line. Three numbers returned a recording: "call cannot be completed as dialed" twice, and "not in service" followed by a fast busy tone once. Two numbers returned a plain busy signal with no recording. | Required by DEV-011 (2). | The three numbers with recordings are out of service for the second phone too, so all nine of their calls stay `dead_number`. The two busy numbers are connected lines (see DEV-016 for how a busy signal is read): their three `dead_number` calls (two Sunday calls and one Wednesday call) get appended correction rows coded `ringout` / NO, noted "blocked", per DEV-011 (2). Both lines also rang on at least one other night, which is consistent. **Disclosed:** the check calls were placed by the study's designer, not the blind caller; each is one call with a recorded or tonal result that leaves little room for judgment. | 3 |
| **DEV-012** | 2026-09-28 | 5.2, 6.1 | **Sunday's Cell C session was run and labeled as Cell A.** On Sun 2026-09-27 the caller completed all 100 calls between 20:00:39 and 21:39:51 PT, but under session W1-A: every row is stamped cell=A, week 1, block_start 2026-09-30 10:00, call_id W1-A-nnn, and the calls followed Cell A's dial order. | The logger listed sessions alphabetically, so W1-A (Wednesday) was first and selected by default, and nothing checked the pick against the date. | **What is intact:** the same 100 communities are dialed in every cell, so the sample is unchanged; the calls happened in the registered Cell C window (Sunday 8 to 10pm); the four automatic backups (25/50/75/100) are exact prefixes of the final file. **What moved:** the within-session dial order (Cell A's permutation was used instead of Cell C's) and the cell label. Randomizing order within a session only spreads time-of-night evenly; one fixed random order is still a random order, so this is disclosed rather than treated as fatal. **Handling:** the original file `instrument-W1-A.csv` (SHA256 89695b13bcbe7a29...) is kept untouched with its four partials. A relabeled copy `instrument-W1-C-relabeled-from-W1-A.csv` sets cell=C, block_start 2026-09-27 20:00, call_id to each community's scheduled W1-C id, and protocol_deviation=DEV-012; every other column, including dial_order as actually dialed and dialed_at_local, is verified unchanged. Analysis uses the relabeled copy. **Prevention (2026-09-28):** the logger now lists sessions by date, preselects today's, refuses a session whose date is not today, and sets aside (after saving to a file) any rows stored under a session that were not logged today. That last part matters: the caller's browser still held these 100 rows under W1-A, which would otherwise have shown Wednesday's real session as already complete. | 100 |
| **DEV-013** | 2026-09-30 | codebook fields, 04-analysis 1.1, DEV-011 (3) | **Four instrument fields were never captured the way the codebook defines them.** Found at analysis, after fielding closed. `rings_to_resolution`, `ivr_menu_depth` and `seconds_in_ivr` are blank on all 300 rows. `answered_call_seconds` holds the whole call, from the start of the call card to the logged result, not the codebook's "greeting to hangup". | The call logger (`call-logger.html`, frozen before call #1) was built to capture the outcome with two questions and a timer. It writes those three fields as blank by design and copies the call timer into `answered_call_seconds`. Nobody compared the logger's output columns against the codebook's definitions before fielding. | **Nothing in the primary outcome, the paired tests or the subgroups uses these fields, so no registered estimate changes.** Three consequences, all disclosed: (1) the ring-rule defense in 04-analysis 1.1 rests on `call_ended_by` alone, which was captured on every call; (2) the DEV-011 exploratory count of voicemail at 0 or 1 rings cannot be computed and is not reported; (3) the burden audit in 04-analysis 1.1 overstates the time a community spent on the phone, because it includes ringing and the caller's logging time. It is reported with that label. The three blank columns stay in the released data file as blank, so a reader can see they were not collected. | 300 |
| **DEV-014** | 2026-09-30 | codebook "Commit at the end of every session" | **The Tuesday Cell B export was not committed to git at the end of its session.** Found at analysis. The Sunday (Cell C) and Wednesday (Cell A) exports were committed; `instrument-W1-B.csv` and its four automatic backups were saved to the study folder on 2026-09-30 at 12:22 PT and left uncommitted. | Missed step. The commit that followed the Cell B save covered only the Cell A subfolder. | The final Cell B export is byte-identical to its own 100-call automatic backup (`instrument-W1-B-partial-100.csv`), every row matches its scheduled `call_id`, `dial_order` and `block_start` in `schedule.csv`, and every dial time falls on Tue 2026-09-29 between 20:01 and 21:40 PT. Those checks show the file is the logger's own output and match the schedule, but git cannot prove when it was saved, so the git evidence for Cell B is weaker than for the other two cells. To be committed before any result is published. | 0 |
| **DEV-015** | 2026-09-30 | 04-analysis 1.2 | **For a count of zero, the lower end of the 95% interval is reported as 0.0%, not the value the registered formula gives.** Found at analysis. Ruled by Ed Brancheau 2026-09-30. | The registered method (Wilson score interval, then the finite population correction applied to the half-width) shrinks the interval toward the Wilson center, which sits above zero when nothing was observed. At a count of zero that pushes the lower end above the observed value: `ai_agent` at 0 of 96 reads 0.0% with an interval of 0.2% to 3.6%. An interval that excludes the observed rate is a known artifact, not a finding. | **Scope:** only estimates with a count of zero, which are `ai_agent` in every cell and the soft `answering_service` share in every cell. The lower end is reported as 0.0% and the upper end is kept exactly as `analyze.js` computes it (3.6% for `ai_agent` at 0 of 96). `analyze.js` itself is not changed, so its raw output still shows the registered lower end, and the published page states the rule next to the number. No estimate with a count above zero is affected. **Direction:** neutral. It widens those intervals by 0.2 points toward zero and changes no comparison. | 0 |
| **DEV-016** | 2026-09-30 | DEV-011 (2), 3.4, codebook, analyze.js input | **Two rules the DEV-011 check needed and did not have.** (1) DEV-011 maps the second-phone result two ways only: out of service stays `dead_number`, rings becomes NO. Two numbers returned a plain busy signal, which is neither. (2) `analyze.js` reads every row as given and has no handling for correction rows, so an appended correction would be counted twice. | Found while applying the check. | **(1) A busy signal from the second phone is read as a connected line, so the call becomes `ringout` / NO, noted "blocked".** Protocol 3.4 defines `dead_number` as invalid, disconnected or not in service, and a busy line is none of those; the codebook itself codes a busy line as NO. Reading busy as dead would have removed three real no-answers from the denominators, which flatters coverage, the exact error DEV-011 was written to stop. **(2) The analysis reads `02-instrument-resolved.csv`,** built from `02-instrument.csv` by letting each correction row replace the row it corrects (300 rows out, from 303 in the log). `02-instrument.csv` keeps every original row plus the three corrections, append-only. `analyze.js` is unchanged. **Effect, both reported:** with the corrections, Sunday is 53.6% (52 of 97) and Wednesday 77.3% (75 of 97). The pre-registered DEV-011 sensitivity analysis, which drops every call noted "blocked", reproduces the uncorrected figures: Sunday 54.7% (52 of 95), Wednesday 78.1% (75 of 96). The Wednesday versus Sunday discordant counts (30 and 7) and the McNemar p-value are identical under both, because each corrected call is a no on both nights of its pair. **Direction:** against the sponsor. The corrections lower the after-hours and weekday rates slightly and leave the gap test unchanged. | 3 |
| **DEV-017** | 2026-09-30 | 04-analysis 5, 5.1, 1.1, DEV-009, protocol 12.2 | **The public release withholds three registered columns, and two other published items are trimmed, because the registered release would let a reader trace individual communities to their own results.** Found at analysis, before anything was published. | 04-analysis 5.1 says the shuffled, coarsened file "cannot" be joined back to the draw. That is not true of the registered columns. The draw is public and reproducible, so anyone can rebuild the list of 100 communities. Among those 100, `region` + `capacity_band` + `publishes_rate` + `rate_band` single out 6 communities, and adding `prospect_overlap` singles out 12 (28 sit in groups of two or fewer). The two rate columns are read off each community's own public website, so a determined reader can recode them. Protocol 12.2 ("never referenced to that community. Ever.") outranks the release list, since the release list was written to serve it. | **(1) Row-level file:** `publishes_rate`, `rate_band` and `prospect_overlap` are withheld. `region` and `capacity_band` stay; with those two, every community sits in a group of at least 3. The three withheld subgroups are published as aggregates only, from `analyze.js` output. One column is added, `noted_blocked` (TRUE for the three DEV-011 corrections), so a reader can rerun the DEV-011 sensitivity analysis. Verified: running `analyze.js` on the released file reproduces every primary, paired, AI-agent, soft-secondary, descriptive and capacity-band figure in `results.json` exactly. **(2) DEV-009 sensitivity:** reported in words, not numbers. The two shared lines are identifiable from the public roster, and any published difference between the main and sensitivity figures would reveal how those specific lines were answered. The numbers are in `results-sens-dev009-shared-line.txt`, which stays local. **(3) Public copy of this log:** community numbers, names, license numbers, phone numbers, call ids, account emails and the caller's name are withheld, and one personal detail is reworded. Everything else is published word for word. | 0 |

---

## Pre-fielding design changes

Not deviations. These predate registration and are recorded here for the audit trail, because the
protocol is registered as v0.1 and a reader should be able to see what moved before it froze.

| Date | Change | Why |
|---|---|---|
| 2026-07-16 | After-hours window moved from 7:00pm to 8:00pm-10:00pm | 7pm licenses neither the control claim nor the "9pm on a Sunday" claim the study exists to support. |
| 2026-07-16 | Added cell B (weekday 8-10pm) | Without it, A vs C confounds time-of-day with day-of-week. |
| 2026-07-16 | n raised from 50 to 100 | At n=50 a 30% rate carries a 19%-44% interval. Not publishable. |
| 2026-07-16 | Private-pay changed from an inclusion filter to a covariate | Filtering would mean ~860 judgment calls by the person who wants a particular answer. |
| 2026-07-16 | Human answering service coded YES, not NO | Coverage measure. The consistency argument is a quality argument and belongs in the secondary. |
| 2026-07-16 | Spacing assertion corrected from elapsed-hours to calendar-days | The elapsed-hours check failed a schedule that satisfied the actual claim (Tue 10pm to Sun 8pm reads as 4.93 hours-days but is 5 calendar days). The check was wrong, not the schedule. |
| 2026-07-16 | `licensee_multisite` renamed `licensee_string_repeated`; subgroup then dropped | The field did not measure what its name claimed. CA per-property LLCs make exact licensee matching report Atria (32 sites) and Pacifica (36) as independent. Ed dropped the cut. See `04-analysis.md` 2.1. |
| 2026-07-16 | Added `ai_agent` outcome | Ed's catch: a conversational AI answering had no code, which under 11.3 would have halted fielding. Coded NO on the primary, reported separately, with a pre-registered "handled by anything" secondary. |
| 2026-07-16 | Cost corrected 10-12h to ~6.75h; per-call slot changed to a 45-min session block | The old schedule chained the rater to a full 2h window to place ~40 min of calls. The window is the randomization space, not the working time. |
| 2026-07-16 | Rulings locked: Q1 CA-only, Q2 no per-call disclosure, Q3 no second rater, Q4 n=100, Q5 subgroup dropped, Q6 git enabled | Ed's rulings. |
| 2026-07-17 | Roster refresh run. Confirmed `file_date` 05252025 is the CURRENT published CDSS RCFE roster (portal `last_modified` 2025-05-27, byte-identical, zero rows changed). Committed snapshot `roster-cdss-rcfe-2025-05-25.csv`. | Verification 11.1. The frame already uses the latest official data; not stale in the sense of a newer version existing. |
| 2026-07-17 | `00-draw.js` guard changed: removed the hardcoded refusal of `file_date` 05252025; now requires `--retrieved-on` and records roster provenance (file_date, retrieved-on, portal-modified, age note) in the manifest. | The old guard would have blocked the real draw on the current-published file. Provenance is the correct control, not a date blocklist. |
| 2026-07-17 | Seed mechanism chosen: public future randomness (first CA SuperLotto Plus draw after registration). | Makes "you redrew until you liked the sample" impossible, not just false. See 06-osf-registration.md A. |
| 2026-07-17 | Caller changed from Ed to a blind VA contractor. | Removes the vendor-bias limitation (author scoring own calls). VA kept blind to sponsor + hypothesis. Notes: offshore voice (cancels in paired contrast), US caller-ID line required. Protocol 6.1, 7.4. |
| 2026-07-17 | Human subtypes replaced with a two-question model. Old `staff`/`service`/`ivr_human`/`deflected` collapse into `human` + two scripted fields: `could_take_inquiry` (Q1) and `human_location` (Q2, asked only when Q1=No). New column `human_location` added; instrument now 21 fields. | Ed's catch: a single caller cannot reliably eyeball on-site-staff vs answering-service mid-call. Asking is objective and scriptable, and it matters now that a VA codes, not the author. Protocol 4.4. Logger, script, codebook, analysis, OSF package all updated and re-verified. |
| 2026-07-17 | Rescoped to Southern California (8-county, frame 428 vs 860) and compressed fieldwork to a single week (3 sessions, C Sun / B Tue / A Wed, all 100 communities each; fixed order; min gap 1 day). Calling on one free Google Voice number (JustCall was $78/mo min; multi-number dropped for the study). | Ed's ICP is SoCal; one-week hits the weekend launch; costs must stay near zero. |
| 2026-07-17 | Seed rule wording corrected before registration: "in official drawn order" -> "sorted in ascending numerical order". calottery.com publishes the five white balls sorted ascending, not in drawn order, so the original wording was not reproducible from the public result (the whole purpose of the seed). Anti-cherry-pick property is unchanged (the numbers still do not exist until the post-registration draw); only the ordering is now unambiguous. | Reproducibility. Caught while building the automated seed fetch. |
| 2026-07-19 | Disclosure on registration visibility: the registration was submitted and frozen 2026-07-17 23:49:30 UTC (OSF `date_registered`), but OSF held it in "pending registration approval" until the sole contributor approved it on 2026-07-19, so it only became PUBLICLY visible on 07-19. OSF's pending state permits approve or cancel only, not editing, so no content changed in the interval, and the authoritative registration timestamp remains 07-17. The ordering the study depends on is unaffected: registration (07-17) precedes the seed draw (07-18 SuperLotto #4100) which precedes call #1 (not yet placed). Recorded because the study's claim is that the rules were public before the data, and the public-visibility date differs from the registration date. | Full transparency; pre-empts the obvious question of why the public timestamp and the visible-since date differ. |
| 2026-07-19 | Quo trial cancelled before it billed; the [number withheld] line is released. A replacement US calling line will be chosen before the 07-26 field week (free Google Voice, or a few dollars of pay-as-you-go VoIP) and logged here once picked. Note the registered protocol attachment names "a single Quo line"; the design requirement that phrase encodes is ONE consistent US number with call recording off, which any replacement satisfies. Recorded here rather than corrected in the protocol, because the registration is frozen (rule 4 above). | Ed declined to start a paid subscription before fielding, and the trial would have expired before the 07-26 field week regardless. |
| 2026-07-19 | Field week pushed one week: field-start Sunday moved 2026-07-19 -> **2026-07-26** (C Sun 07-26, B Tue 07-28, A Wed 07-29). Sample and per-session dial order are UNCHANGED: the same registered seed (0203102737) was re-used and the community_id -> facility_number mapping verified identical before and after. Zero calls had been placed, so this is a pre-fielding schedule change, not a protocol deviation. | No blind caller was available for the 07-19 week (both contracted VAs went unresponsive). Fielding with the study's own designer as the caller would have contradicted the blind-caller design that is registered on OSF and published on the public method page (protocol 7.4), and would have reinstated the exact vendor-bias attack that design removes. Moving the week preserves the design at a cost of one week. |
| 2026-09-25 | Logger: every result now needs a confirm click, and the confirm button and the next card ignore input for 0.7 seconds. | The Friday practice run showed a likely accidental double-click advancing a call. Interface only; no code, definition or outcome changed; zero real calls placed. |
| 2026-07-17 | Calling tool finalized: **Quo** (formerly OpenPhone) Starter free trial, account [account withheld], US number [number withheld], auto call recording OFF (CA two-party consent). Supersedes the "Google Voice" line named in the row above. VA dials via shared login; trial cancels ~2026-07-23. | Quo's trial gives a clean US caller-ID line at zero cost for the study window, and recording-off matches the no-audio design. |
