CASE STUDY · PUBLIC DATA · AI PIPELINE · 2026

We read fifty years of federal election law with an API and a prompt.

Every FEC advisory opinion since 1975 — more than two thousand PDFs — pulled from the government's public OpenFEC API, run through an AI extraction pipeline, and turned into a searchable, standardized digest. Live at fecadvisor.com.

0+
Opinions processed
0+
Years covered
0
Fields per opinion
[X] hrs
Full-corpus runtime

▸ each point above is one opinion. scroll to run the pipeline.▌

What it is

FEC Advisor: a public research tool that summarizes every Federal Election Commission advisory opinion in a fixed format and makes the whole body of law searchable.

What we did

Designed and built the full stack: OpenFEC ingestion, document text extraction, AI summarization against a fixed schema, structured storage, search, per-opinion pages, export, and automated ingestion of new opinions.

Context

Commissioned by a campaign finance law practice to finish a research project their clerks had started by hand years earlier; funded by the Coolidge Reagan Foundation.

The problem

Two thousand PDFs and no index

Advisory opinions are how the FEC tells the public what campaign finance law means in practice. Anyone can ask the Commission whether a specific plan is legal, and it answers in writing. There are more than two thousand of these opinions going back to 1975. They're all public. They're also all PDFs, with no consistent structure and no way to search across them by what was asked or how it was decided.

01

A manual project that couldn't finish.

Years earlier, a law firm had put clerks on this: read one opinion, write a standardized summary — who asked, what they asked, what the Commission decided — move to the next. The summaries were useful. The project covered a fraction of the corpus before it stalled, and the Commission kept issuing new opinions.

02

Unstructured source material.

The FEC's public data is thorough but raw: document listings, PDFs, scanned pages from the 1970s and 80s. Nothing about it is ready to query.

03

No consistency across summarizers.

Different clerks wrote differently. Without a fixed schema you can't compare outcomes across decades, filter by statute, or spot how reasoning shifted.

04

Cost per document.

At clerk-hours per opinion, completion was a multi-year commitment nobody could justify. So it stayed unfinished.

[Quote about the original clerk project — e.g. "We got a few hundred done over several years. It was some of the most useful research we had, and we could never finish it."]
— [Name, Title]
The pipeline

The clerk project as software

We rebuilt the manual process as a pipeline. The OpenFEC API lists every opinion and its documents. We pull the text, run each opinion through an AI model with a fixed extraction prompt, store the answers as structured records, and put a search interface on top. New opinions are picked up automatically. The whole corpus can be re-processed in hours if the prompt improves.

terminal
$ ▌
01 · FETCH

Fetch

We start with the FEC's own public API. It knows every opinion the Commission has ever issued and every document attached to it.

02 · EXTRACT

Extract text

Each opinion is a PDF, some of them scans from the 1970s. We pull the text out and clean it.

03 · AI

AI extraction

One fixed prompt asks the same six questions of every opinion. Same questions in 1975, same questions today.

04 · STRUCTURE

Structure

The answers are validated and stored as one JSON record per opinion. Every opinion, same shape — that's what makes the corpus comparable.

05 · SHIP

Index & ship

Search, browse by year, a stable page for every opinion, CSV export — and a scheduler that picks up new opinions as the Commission publishes them.

What shipped

What Shipped

Full-corpus ingestion

Every advisory opinion from 1975 to present, discovered and fetched through the OpenFEC API.

FETCH
Fixed six-field extraction

Requestor, question presented, outcome, Commission vote, key legal issues, statutes and regulations cited. Same fields for every opinion, every year.

EXTRACT
Structured storage

One JSON record per opinion, queryable and re-processable in bulk.

STRUCTURE
Browse by year

1975 to the current year.

INDEX
Search

Keyword and AO-number search across the digest.

INDEX
Per-opinion pages

A stable URL for every opinion (e.g. /opinion/2006-02) with its summary and links to the original documents on fec.gov.

SHIP
Source-linked by design

Every summary points at the authoritative opinion. The tool is an index to the law, not a substitute for it.

SHIP
CSV export

Batches of summaries download for briefs, memos and research spreadsheets.

SHIP
Automated updates

Scheduled checks against the API pick up newly published opinions and process them without anyone touching it.

FETCH
The corpus, again

The full corpus. Every point covered.

New opinions are processed as they're published.

Try it

Try the digest

[Insert 6 real summaries from fecadvisor.com]

This is a static sample. The live tool at fecadvisor.com runs the same interface over the full corpus.

Before / after

Same opinion, two ways

[Screenshot: the same opinion on fecadvisor.com]

[Screenshot: original AO PDF page from fec.gov]

⇔

Same opinion. The left takes ten minutes to read. The right takes ten seconds, and links back to the left.

Results

What changed

0%

of published advisory opinions summarized [confirm]

[X]%

of the corpus the manual project had covered when it stalled

[$X]

total AI processing cost for the entire corpus

[X] min

to find every opinion citing a given regulation, down from [Y]

Manual (clerk hours)
[X]
Pipeline (API cost)
[Y]

Finished.

Every opinion the Commission has issued has a summary in the same format, from 1975 to the most recent. A project that had been "someday" for years is a completed dataset.

Stays finished.

New opinions are processed as they're published. The digest doesn't decay the way the manual archive did.

Searchable in ways PDFs never were.

Every opinion citing a specific statute, every request from a particular kind of entity, every outcome on a recurring question: a search, not a reading assignment, with every result linking to the authoritative source.

[Quote about using the finished tool]
— [Name, Title]
What it is and isn't

FEC Advisor is a research aid, not legal advice. Every summary is generated by an AI model from the Commission's own published documents, and every summary links to the original so it can be verified in one click. The tool tells you which opinions to read and what they said in outline; the authoritative text is always the FEC's.

We built the pipeline to be re-runnable. If the extraction prompt gets better, or the schema gains a field, the full corpus re-processes in hours. That's the difference between a dataset and an archive.

Delivery

How We Delivered

  1. #01 · [Mon YYYY]

    Proof of concept in a day

    A working prototype against the live OpenFEC API, generating summaries on demand, to validate output quality before committing to the full build.

  2. #02 · [Mon YYYY]

    Extraction schema

    Locked the six fields and the prompt, tested against the clerks' original hand-written summaries as a quality benchmark.

  3. #03 · [Mon YYYY]

    Full-corpus run

    Batch processing of 2,000+ opinions with retries, rate-limit handling and validation on every record.

  4. #04 · [Mon YYYY]

    Product

    Search, year browse, per-opinion pages, CSV export, public site.

  5. #05 · [Mon YYYY]

    Automation

    Scheduled ingestion of new opinions.

  6. #06 · [Mon YYYY]

    Launch

    fecadvisor.com.

Tech

Tech & Details

source
FEC OpenFEC API (public, official)
extraction
Anthropic Claude, fixed prompt, validated JSON out
storage
One structured record per opinion; bulk re-processable
interface
Web app: year browse, keyword/AO-number search, stable per-opinion URLs
export
CSV
updates
Automated ingestion of newly published opinions
disclaimer
AI-generated, source-linked, not legal advice

Sitting on a public dataset or an archive nobody has time to read?

We turn government APIs and unread documents into tools people use every day.