Guides · AI & clinical paperwork

Should you use ChatGPT to fill out FMLA paperwork?

TL;DR

A chatbot can draft. It can't certify. General chatbots are genuinely useful for understanding the WH-380 forms — but for producing the certification itself they have four structural gaps: they can invent dates and numbers that were never in your notes, they output text rather than the official PDF itself, they improvise FMLA's regulatory rules instead of applying them, and consumer tiers give PHI no safe place to be (typically no BAA, variable retention). If you do use one, verify every number against your own records before anything reaches a form you sign.

It's a fair question — and clinicians are already doing it. Describe the patient to a chatbot, ask for the WH-380-E answers, copy them onto the form. Sometimes that works. This article is an honest accounting of when it does, when it fails, and what to check before you sign anything a language model produced. (Yes, we make purpose-built software for this — which is exactly why we know where the failure modes are. We test language models against these forms constantly.)

What chatbots are genuinely good at here

  • Explaining the form. "What does 'incapacity plus treatment' require?" gets a solid answer with citations you can check.
  • Drafting language. Turning "she can't lift her arm above shoulder height" into a clean Part C sentence is squarely in a model's wheelhouse.
  • Organizing a messy case. Pasting a timeline and asking for a summary of dates and visits is fast and usually accurate — when every date in the output already existed in the input.

Where they fail: four structural gaps

1. Invented specifics — the date problem

Language models produce plausible text, and plausible is exactly the failure mode a certification cannot tolerate. In our own testing of frontier models against these forms, a model given "neurology follow-up monthly from 07/15/2026 through 10/15/2026" expanded it into a list of seven specific appointment dates — none of which existed anywhere in the source. Each looked completely reasonable. A model told a patient was "off work about three weeks after discharge" will often volunteer an exact return date it computed itself. In a chat window, nothing stands between that invented specificity and your signature except your own proofreading — of every single number, every time.

2. Text out, not the form out

The certification isn't prose; it's a specific Department of Labor PDF with checkboxes, phrase-choice items ("has been / is expected to be"), and a signature block. A chatbot gives you text to re-type — re-introducing transcription errors — or an attempted recreation of the form that isn't the official document. The DOL PDFs are also technically odd under the hood (non-standard checkbox groups, internal values that differ from the printed text), which is why "just fill the PDF" is harder than it sounds even for software, let alone a chat window.

3. Improvised rules instead of enforced ones

The WH-380 forms encode regulatory judgments that general models get wrong in predictable ways. Three we see repeatedly in testing:

  • The >3-consecutive-days test: a patient who misses one day per migraine, three times a week, does not meet "incapacity plus treatment" — that's a chronic condition with intermittent leave. Models pattern-match "misses lots of work" into the wrong category.
  • Relationship categories on the family-member form (WH-380-F): a severely ill 15-year-old is always "child under 18" — never "adult child incapable of self-care," no matter how much care they need. Models over-weight severity and pick the wrong box.
  • Who signs what: Section II of the WH-380-F is completed and signed by the employee, not the provider. A chatbot filling "the whole form" doesn't know where your authority stops.

An insufficient or wrongly-completed certification isn't cosmetic: the employee gets a cure notice against a deadline, and the leave request can ultimately be denied.

4. PHI has nowhere safe to be

Pasting a patient's situation into a consumer chatbot makes you the compliance officer for that decision. Consumer tiers generally do not offer a HIPAA Business Associate Agreement, and retention and training policies vary by product and setting. Before any patient details go into any AI tool, you should be able to answer: Does the vendor sign a BAA? Is the conversation retained? Can it be used for training? Who can see it? If the answers aren't clearly yes-BAA, no-retention, no-training — de-identify ruthlessly or don't paste.

If you use a chatbot anyway: the checklist

  1. De-identify. No names, no MRNs, no employer names — the model doesn't need them.
  2. Verify every number. Each date, frequency, and duration in the output must exist in your own records. Treat any date you didn't provide as invented.
  3. Check the category logic yourself against the definitions printed on page 4 of the form (or our field-by-field guide).
  4. You transcribe, you sign, you own it. The chatbot's fluency doesn't transfer any responsibility.

What purpose-built software does differently

PatientPapers was built around exactly these four gaps, so the fixes are structural rather than procedural:

  • Nothing is invented, because nothing is generated: there is no language model in the product. Guided questions are how the form gets completed. If you paste a note instead, a deterministic parser proposes values and shows each one beside the words it came from — and nothing reaches the form until you confirm it, one value at a time. All arithmetic — leave windows, verb tenses, durations — is code.
  • The genuine form: output is the actual DOL PDF with your confirmed values on it, signature applied, flattened — not text to re-type.
  • Encoded rules: the >3-day test, the relationship categories and who signs Section II are checked by the software, and GINA and diagnosis-disclosure problems are flagged before you sign. They flag; you decide. Nothing here relieves you of reading the form.
  • One patient-data path, and it is empty: nothing you enter is sent anywhere. The app asks our servers for the blank forms and field help it needs, and sends nothing back — so there is no patient-data vendor to vet, no patient-data retention policy to read and no BAA question to answer, because no patient information ever reaches us.

FAQ

Can ChatGPT fill out a WH-380-E for me?

It can draft answers as text that you re-type onto the official PDF. It can't produce the official form itself, and it may invent dates or pick the wrong regulatory category — every number needs manual verification before you sign.

Is it safe to paste patient information into a chatbot?

Consumer tiers generally offer no BAA, and retention/training policies vary. Know the answers to BAA / retention / training questions before any patient details go in — or de-identify completely.

What's the single biggest risk?

Invented specifics: plausible dates, frequencies, and durations that were never in your source material. On a signed federal certification, those become your problem, not the chatbot's.

How is purpose-built software different?

Every value is confirmed by you before it reaches the form, the output is the genuine DOL PDF rather than text to re-type, FMLA rules are checked by code, and nothing you enter is sent anywhere.

Disclaimer: This article is general information, not medical, legal, or compliance advice. Statements about consumer AI products describe common characteristics of consumer tiers as of the publication date; check any vendor's current terms directly. You are the certifying provider and are responsible for the accuracy of everything you certify.

Answer. Confirm. Sign.

PatientPapers puts your confirmed values on the genuine DOL certification — WH-380-E and WH-380-F — with every value shown for your confirmation before it lands. Nothing you enter is sent anywhere. $25 per month, per certifying provider.

Join the mailing list