Method

How we know this

This page is the project's audit trail. It describes what the system actually does — not what it was designed to do — and states the limits plainly, because a research method you can't check isn't worth much.

1. The question, and why it takes two kinds of data

"What is this career actually like?" has two halves, and each available answer covers only one of them. Government statistics tell you how much the job pays, how many people do it, and where it's growing — measured properly, on representative samples, with denominators you can trust. They will never tell you what the fourth night shift in a row feels like, or which part of the job nobody warns you about.

First-person accounts tell you exactly that, and are useless for measuring how common anything is — the people who write them chose to. So this site carries both, side by side, and tries very hard never to let one masquerade as the other.

2. Where the data comes from

Two numbers matter here and they are very different: what has been collected, and what actually reaches the page. Both are shown, measured by counting the pipeline's own outputs.

What you're actually reading

40 displayed quotes
  • Reddit forumforum
    37 · 92.5%
  • Federal rulemaking commentson the record
    3 · 7.5%

What we've collected

404,105 rows
  • Reddit forumforum
    310,539 · 76.8%
  • YouTube comments
    47,616 · 11.8%
  • Professional forum (SDN)forum
    40,377 · 10%
  • Federal rulemaking commentson the record
    1,924 · 0.5%
  • Podcast transcriptspublished
    1,545 · 0.4%
  • Stack Exchangeforum
    1,219 · 0.3%
  • Nursing blogspublished
    528 · 0.1%
  • Web pages (Common Crawl)published
    289 · 0.1%
  • Congressional testimonyon the record
    63 · 0%
  • Personal essayspublished
    5 · 0%

Collected ≠ displayed: most collected rows never pass screening, and some sources have produced no displayed quote at all yet. Both numbers are shown so the concentration is visible rather than hidden. Measured 2026-08-15 by counting the pipeline's own outputs; nothing here is estimated.

What each source is worth

  • Federal rulemaking comments — submitted by name to a federal agency on a public docket (for example the CMS minimum-staffing rule). Filing false information with a federal agency is an offence, and these are written to be read by regulators. They are the most accountable material in the corpus, which is why they are badged distinctly wherever they appear.
  • Congressional testimony — statements entered into the congressional record. 63 documents are collected; none has yet produced a displayed quote, so you are not currently reading any.
  • Public forums (Reddit, Student Doctor Network, Stack Exchange) — anonymous, candid, and self-selected. Reddit is the large majority of what is displayed. Nobody here is accountable for what they say, which cuts both ways: less posturing, and no consequence for exaggeration.
  • YouTube comments, blogs, podcasts, essays — public, anonymized, same self-selection caveat as any open platform.

3. How content is screened

Every piece of text passes three gates before anything it says can appear:

  1. A keyword gate. The text must contain a career-specific term — for nursing: nurs*, RN, BSN, ADN, NCLEX, CRNA, LPN, LVN, NP. This runs on the quoted sentence itself, not just the surrounding post, so a model can't justify an off-topic quote with on-topic reasoning.
  2. Two independent models must agree. One model classifies the item as relevant; a second opinion is then requested from a different provider, explicitly excluding the one that already answered. If no other provider currently has quota, the pipeline waits rather than accepting a single model's word — the fallback is delay, never a lowered standard.
  3. Evidence-verified extraction. Each answer must quote a verbatim sentence from the source. The quote is checked against the source text mechanically; if it isn't found character-for-character, the answer is discarded.

What that rejects, in real numbers

StageRowsSurvives
Collected from forums (Reddit corpus)310,539
Put through model classification44,01214.2%
Produced any usable extracted answer2570.58%
Two-model confirmed (tier A)1880.43%
Passed the answerhood check and are displayed400.091%

Read that bottom row honestly: of 44,012 conversations put through classification, 40 produced a quote good enough to publish. The pipeline throws away almost everything it sees. Counts as of 2026-08-15.

The answerhood check

A quote can be genuine, verbatim, and still not answer the question it's filed under — a commenter arguing with someone else, or advice aimed at one person. So displayed quotes face a fourth gate: two independent models must both agree the quote actually answers its question. That check most recently cut the pool from 141 candidates to 45. Nine quotes the models split on were promoted by hand after review; each is recorded individually with its reason, and the rest stay out.

4. What is not verified

Screened means two independent AI models had to agree the account is genuinely about this career, and every extracted answer must quote a verbatim sentence from the source or it is dropped. We do not verify anyone's identity or employment.

Nobody uploads a licence. If a person writes convincingly about nursing and two models agree the content is genuinely about the career, their words can appear here. That is a real limit, and the site says "screened", never "verified", because of it.

If something of yours is on this site and you want it gone, ask — no account, no explanation required.

5. Known biases and limits

  • Self-selection. People post when something is worth saying — usually when it's wrong. Contentment is quiet.
  • Platform concentration. One source dominates what you read (see the mix above). Any quirk of that platform's culture is baked into the corpus, and this is the single biggest weakness of the project.
  • Survivorship, in both directions. People who left the profession and people who stayed post for different reasons; neither is weighted, because there is no honest way to weight them.
  • Identity is unverified. See above.
  • Frequency is not prevalence. A theme appearing often here means it is said often in this corpus. It does not mean that share of nurses experience it. When you want "how common", the survey data answers that; the accounts answer "what specifically".

6. What we refuse to do

Three sources are documented as blocked rather than scraped around. They are named here because a method section that only lists successes is advertising.

  • allnurses (forums)Cloudflare bot-protection on all paths, including robots.txt
  • Goodreads (reviews)robots.txt disallows /book/reviews/, /review/show, /review/list
  • Amazon (reviews)anti-bot 503 on polite fetch (robots permitted, edge refused)

Being blocked ends the matter: these domains are also excluded from the Common Crawl miner, so a bot wall is not routed around via archives or caches. robots.txt and site terms are respected, requests are rate-limited and identify themselves, and no login-walled or private content is touched.

7. Official statistics used

DatasetVintageUsed for
BLS OEWSMay 2024 (state file); May 2023 percentilesWages by state, wage percentiles, ladder occupations
BLS Current Population Survey2000–2025 annualThe 26-year pay history and the gender pay gap
BLS Employment Projections2024 base → 2034Employment now vs projected, annual openings
BLS SOII2023–2024Injury and illness cases with days away from work
O*NETdb 30.0Core tasks, task importance and frequency, work activities
College Scorecardfield-of-study, latest releaseTuition, graduate earnings and debt by credential
NCES Digest table 330.102023 edition (1963–2022)Cost-of-college backdrop for the indexed timeline
HRSA Workforce ProjectionsFY2025Share of 2038 demand met, by state
Census population estimatesNST-EST2024RNs per 1,000 residents
BEA Regional Price Paritieslatest releaseCost-of-living adjustment behind 'real pay'
GSS / NIOSH Quality of Worklifemodule last fielded 2014The representative satisfaction anchor (n=42 nurses)
State boards of nursing (NCLEX)2021–2025 by stateFirst-time pass rates per program
CCNE / ACENcurrent directoriesProgrammatic accreditation status

8. How often this updates

Collection runs daily and continuously adds to the corpus. A separate daily job re-exports every dataset from the pipeline's outputs, rebuilds the site and republishes it, so the numbers on every page — including this one — move as the corpus grows. Any section without enough data to be honest renders an empty state explaining what is missing, rather than filler; the satisfaction-over-time chart is the current example, held back at 24/50 scores.

Data on this page generated 2026-08-15.

9. What this site collects about you

Page views are counted with GoatCounter: no cookies, no cross-site tracking, no personal data, and nothing that identifies you. It records which pages are viewed and roughly where in the world from — that is all, and it is the only analytics on the site. There is no Google Analytics, no advertising pixel, and no third party receives anything about you.

If you use the "was this useful?" buttons, we store your yes/no, any line you type, and which page you were on. If you fill in the survey, we store your answers and — only if you choose to give it — an email address, kept apart from the answers. Nothing else. You can ask for any of it to be removed.