Home / Case study
Case study · Tax & information reporting · Digital transformation
From December chase emails to a 1099 readiness workflow.
It arrived as a request for a better W-9 tracker. The real problem was three systems creating the same vendors twice, so the chase list was wrong before anyone opened it. The fix turned a manual season into an automated readiness process, with people still making every tax decision.
From our founders’ earlier work, before MizDauntless. Company details removed; figures rounded.
- Organization
- Mid-market software company
- Vendors
- ~4,200 · ~2,600 reportable
- Forms
- 1099-NEC and 1099-MISC
- Teams
- AP, tax, payroll, controller, filing provider
Layer 1 · Outcome
One season later.
Layer 2 · Workflow
Follow the work, lane by lane.
Manual or reworkAutomatedA person decides
Show the full workflowHide the full workflow
Tap any step for detail.
01Three systems create vendors separatelySystemsDuplicates
The ERP, a contractor-payment tool and an expense platform each add vendor records. About 6% of reportable vendors exist twice under slightly different names.
02Export and filter by handAccounts payableManual
AP exports the vendor master and filters to anyone paid over $600. The list is wrong before anyone opens it.
03Chase W-9s by emailAccounts payableRework
Six people, one spreadsheet. Duplicate vendors get two emails and usually ignore the second.
04Send W-9s in any formatVendors
Fillable PDFs, scans, phone photos and the occasional wrong form, all by email.
05Log documents in the spreadsheetAccounts payableRe-entry
When the W-9 lands on one duplicate and the payments on the other, it looks like the vendor never sent it.
06Check every exceptionTax teamBottleneck
TINs and names checked by hand. AP can’t read the failure messages, so every exception waits for a tax analyst.
07Build the filing fileTax team15 days
Fifteen days from the AP extract to a file the provider accepts.
08File, then correctFiling providerRework
210 corrections after submission, and about 190 IRS name/TIN mismatch notices (B-notices) the following year.
01Match vendors across all three systemsSystemsAutomated
Likely duplicates are ranked for review, including near-name matches that exact matching misses.
02Merge duplicatesAccounts payablePerson decides
A person makes every merge decision and records why.
03Work out who needs a W-9SystemsAutomated
Required documents are computed from payments and vendor type, so nobody audits the vendor master by hand.
04Send drafted W-9 requestsAccounts payablePerson sends
Each request is pre-filled with what’s missing and why. A person reviews and sends it.
05Return W-9sVendorsMostly automated
About two-thirds are read straight from the form’s fields. The rest are read automatically, then checked by a person.
06Check before filingSystemsAutomated
TIN format and name/TIN match checked on every record, with a trail back to the source row.
07Close routine exceptionsAccounts payable65% self-serve
Plain-language explanations let AP fix missing forms and malformed TINs without waiting for tax.
08Decide the judgment callsTax teamPerson decides
Classification and anything uncertain stays with tax, and every approval is named and traceable.
09Receive a filing-ready fileFiling provider4 days
Four days from extract instead of fifteen.
Layer 3 · How it was built
Five stages.
01Discovery and baselineSix weeks, four functions interviewed separately, documents over opinions.
AP, tax, payroll and the filing provider were interviewed separately before meeting together, and the work started from the prior-year filing file, the correction file and the exception tracker rather than from descriptions of the process.
The turning point came walking the AP manager through one real case. Asked what happens when one duplicate record holds the W-9 and the other holds the payments, she paused: “Then it looks like they never sent it.” The problem wasn’t the tracker. It was vendor identity.
The baseline was then measured from system and email timestamps, because the first week of self-reported effort estimates proved unreliable.
02Decide what automation may never doA one-page boundary, signed by the tax manager and the controller.
Clear, versioned rules make every tax determination. Automation and AI may suggest, rank, extract, explain and draft; they may never decide reportability, merge records or approve anything. A person approves.
That boundary is what made the system trustworthy enough to deploy, and it’s why AP could later act on explanations without routing to tax.
03Clean the foundation firstResolve vendor identity before building anything that counts.
Vendor records from the three systems were matched and duplicates resolved, because every downstream number depended on a correct population. Every imported value traces back to its source file and row. Tax IDs are masked on screen and matched by secure hash, and each team sees only what its role needs.
04Automate the checksCheapest reliable method first; people keep the judgment calls.
W-9s go through a ladder: read the form’s fields directly, then the text layer, then automated reading, then a person. About two-thirds never needed anything beyond the first steps. Every extracted value passes the same validation as imported data, and tax ID and classification are always reviewed by a person.
Exceptions arrive with plain-language explanations and chase emails are drafted for a person to send. An auto-approve feature was built, tested and switched off: approval is the control, and automating it deletes the control.
05Roll out, adopt and measureA three-week parallel run, then an adoption stall and its fix.
The new path ran alongside the manual one for three weeks and results were compared. Adoption still stalled at about 30% in the first month. What turned it round was showing the tax analysts the log of cases where the system declined to guess: trust came from seeing it hold back.
AP initially read the change as extra work. Measuring what AP cared about, how fast its own exceptions closed, made the case: AP stopped waiting on tax.
Layer 4 · Results
What changed, and what drove it.
Most of the improvement came from clear rules and clean data, not AI. AI earned its place in two spots: finding near-duplicate vendors, and explaining exceptions so AP could close them.
Year-end preparation effort (tax + AP)~480 hrs~190 hrs
What drove it: Rules and AI together. Three kinds of work stopped existing: chasing duplicate vendors, auditing the vendor master by hand, and routing every exception through tax.
How it was measured: Tax and AP time across the November–January season, sampled from queue and email timestamps after self-reported estimates proved unreliable.
Days from AP extract to filing-ready file154
What drove it: Rules and AI together. Clean vendor identity and automatic checks meant the file was right the first time instead of after rounds of fixes.
How it was measured: Time from the first import of the AP extract to generation of the filing package.
Reportable vendors missing a valid W-9 (Nov 1)780 · 19%210 · 5%
What drove it: Rules and AI together. Documents were requested from the right vendor record, with drafted requests saying exactly what was missing.
How it was measured: The document-status field on the vendor master at the November 1 checkpoint.
Missing or malformed tax IDs31040
What drove it: Rules. Format checks run on every record at import, so bad tax IDs are caught months before filing.
How it was measured: A format rule applied to every record at import.
Name/TIN match rate before filing88%98%
What drove it: Rules. Names and tax IDs are matched on every record before filing, and mismatches go to a named owner.
How it was measured: The match result stored for each vendor before submission.
Duplicate vendors in the reportable population~6% · 156<1% · 22
What drove it: AI-ranked, person-decided. AI ranked likely duplicates, including near-name matches that exact matching misses. A person made every merge decision.
How it was measured: Duplicate detection at intake, across all three source systems.
Corrections filed after submission21038
What drove it: Rules and AI together. Fewer identity and tax ID errors reached the filing in the first place.
How it was measured: The filing provider’s correction file, reconciled back to source records.
IRS name/TIN notices the following year~190~35
What drove it: Rules and AI together. This is the least direct measure: it lags by a year and other changes are mixed in. The earlier rise in the name/TIN match rate is the leading sign that the two are connected.
How it was measured: IRS mismatch notice (B-notice) volume in the following season.
Exceptions AP closed without routing to tax0%65%
What drove it: AI explanations. Plain-language explanations let AP fix routine exceptions itself. Classification and anything uncertain still went to tax.
How it was measured: The owner recorded on each closed exception.
Tap any row for what drove it and how it was measured.
What we’d do differently
- Build the measurement harness a release earlier.
- Bring compliance into design, not just review.
- Measure accuracy field by field from day one, not as an average.
What it shows
- The stated problem is often a symptom. Evidence comes before renaming it.
- Fix the data foundation before automating anything that counts.
- Automation earns trust when it’s clear what it will never decide.
Run a process on spreadsheets and chasing?
Start with a two-minute brief. If a diagnostic is the right next step, we’ll scope it in writing first.