The raw material

Data explorer

A small, heavily-screened set of first-person accounts — not a searchable database of everything nurses have said. Almost everything collected is discarded before it reaches this page; what survives is quoted exactly as written, with its source attached.

How 40 quotes come from 404,105 rows

Almost everything collected is thrown away. That is the point — but it means the quotes you can browse are a deliberately small, heavily-screened set, not a searchable database of everything nurses have said.

  1. Collected across all sources 404,105
    full corpus
  2. Reddit rows collected 310,539
    full corpus
  3. Put through model classification 44,012
    full corpus
  4. Produced a citable extraction 257
    full corpus
  5. Two-model confirmed 188
    full corpus
  6. Quoted on this site 40
    displayed subset

266,527 Reddit rows (86%) have not been classified yet — collection outruns screening, and the pipeline is still working through the backlog. Today's small quote pool is partly a queue, not a verdict.

Reading these numbers: everything above the last row counts the full corpus and drives the theme rankings, sentiment mix and skew analysis elsewhere on the site. Only the final row is the displayed subset — the quotes you can read. Never treat the two as the same number.

What do the source badges mean?
  • Federal rulemaking comments Filed by name on the public record with a federal agency (e.g. the CMS minimum-staffing rule). Submitting false information is a federal offence, so these carry the most weight of anything here.
  • Congressional testimony Statements entered into the congressional record. Collected, but none currently appear in the displayed quotes.
  • Reddit forum Public posts and comments, anonymized. The largest share of what is shown — candid, but self-selected: people post when something is worth complaining about.
  • YouTube comments Public comments on nursing videos. Collected and used in the aggregate analysis, but not quoted verbatim on the site — so you will not see one of these as a quote.

Screened means two independent AI models had to agree the account is genuinely about this career, and every extracted answer must quote a verbatim sentence from the source or it is dropped. We do not verify anyone's identity or employment.

Where this comes from

The corpus, honestly

One source dominates what you're reading. Showing that is the point: a corpus is only as good as a reader's ability to see its shape.

What you're actually reading

40 displayed quotes
  • Reddit forumforum
    37 · 92.5%
  • Federal rulemaking commentson the record
    3 · 7.5%

What we've collected

404,105 rows
  • Reddit forumforum
    310,539 · 76.8%
  • YouTube comments
    47,616 · 11.8%
  • Professional forum (SDN)forum
    40,377 · 10%
  • Federal rulemaking commentson the record
    1,924 · 0.5%
  • Podcast transcriptspublished
    1,545 · 0.4%
  • Stack Exchangeforum
    1,219 · 0.3%
  • Nursing blogspublished
    528 · 0.1%
  • Web pages (Common Crawl)published
    289 · 0.1%
  • Congressional testimonyon the record
    63 · 0%
  • Personal essayspublished
    5 · 0%

Collected ≠ displayed: most collected rows never pass screening, and some sources have produced no displayed quote at all yet. Both numbers are shown so the concentration is visible rather than hidden. Measured 2026-08-15 by counting the pipeline's own outputs; nothing here is estimated.