← Personal projects Data engineering Data engineering project · 2025
Attorney data extraction at scale
Every law firm publishes attorney bios differently, and none of them publish them cleanly. I built a Selenium scraper suite covering forty-plus U.S. firms, normalizing all of them into a single twenty-four-field schema, feeding an AWS ETL job that lands clean records on a schedule.
- 40+
- Firm scrapers shipped
- 24
- Fields normalized
- Manual → scheduled
- Collection cadence
Architecture
- 01Firm siteSelenium, headless Chrome
- 02Profile parserper-firm selectors
- 03Normalizer24-field schema
- 04Validatorsstates, honors, years
- 05Excel + S3openpyxl, AWS ETL
The problem
A sales team needs structured records for practicing attorneys in the United States: where they work, what they practice, where they went to school, which state bars they’re admitted to. That data is public. It’s on every firm’s website. It is also, from a data engineering standpoint, a worst case.
There is no standard. One firm renders bios as server-side HTML with clean semantic markup; the
next builds them client-side from a JSON blob; the next hides half the biography behind a “Read
More” toggle that only fires on click. Education appears as prose in one place and as a list in
another. Some firms put “J.D., cum laude” in a single string; some split degree, school,
honors, and year across four sibling elements; some use an <em> tag for honors and nothing
else.
And the target wasn’t “get the data” — it was get the data in exactly one shape, because everything downstream assumed that shape.
The constraint that shaped everything
The output schema was fixed and non-negotiable: twenty-four fields, in a specified order, covering identity, firm and role, contact, biography, practice areas, two separate education records, bar admissions, languages, and source URLs. Every scraper, regardless of what the source site looked like, had to emit exactly that.
That constraint is the whole story of this project. It meant the per-firm work could only ever be extraction — the moment a scraper started making its own decisions about formatting, the dataset fractured.
So the architecture split hard along that line:
- Per-firm extractors knew about selectors, pagination, and the specific ways one site was weird. Nothing else.
- A shared normalization layer owned every formatting decision, and every scraper called it.
Normalization rules worth naming
Most of the real engineering was here, and most of it came from watching the data fail.
Education had to be split before it was parsed. Law school and undergraduate records are
separate fields, and firms interleave them freely. Parsing had to identify which entry was
which before touching either — and the school-name fields had to contain only the school
name. Not “B.S., George Mason University,” not “George Mason University (2014).” Just
George Mason University. Degrees, majors, years, honors, and punctuation all stripped out
into their own fields or dropped.
Honors were matched, never inferred. Firms write honors in free text, and free text
produces garbage categories. Honors were extracted from a specific nested element and then
matched case-insensitively against a fixed validated list — cum laude, magna cum laude,
summa cum laude, Order of the Coif, Law Review, clerkship, and a small set of secondary
degrees. Anything that didn’t match the list didn’t become a value; it became UNASSIGNED.
A known-unknown is a usable record. An invented category is a corrupted one.
Bar admissions were restricted to U.S. states. Firm bio pages list courts alongside states — district courts, circuit courts, the Supreme Court. Only state-level admissions (plus the District of Columbia) belong in that field, so courts were filtered out, and admission years were stripped from the values rather than left to pollute the string.
Middle names lost their punctuation. Trivial-sounding, and it mattered: J. and J are
different keys, and deduplication downstream ran on name fields.
Biographies had to be complete. Where a “Read More” control existed, the scraper expanded it before reading. A truncated bio isn’t a shorter record, it’s a wrong one.
Some profiles were out of scope entirely. Non-U.S. offices — London in particular — were excluded at the extractor level rather than filtered later, so they never entered the pipeline.
Making forty scrapers maintainable
Forty independently-written scrapers is forty independent liabilities. The suite converged on a single template that every new firm started from:
- A data dictionary initialized with all twenty-four keys, so every record was structurally complete before extraction began — missing fields came out empty, never absent.
- A standard education-parsing block, calling the shared normalizer.
- Per-field
try/exceptisolation.
That third one is the reason the suite scaled. A firm that redesigns its bio template breaks one selector, not one scraper. The record still lands, with one empty field and a logged failure, instead of the run dying at attorney fourteen of six hundred.
What it fed
Output was written through openpyxl into the fixed template structure, then picked up by an
AWS-based ETL job that loaded it into the sales database. The manual collection process it
replaced was a person opening bio pages and copying fields into a spreadsheet.
What I’d do differently
Two things.
Contract tests per firm. The try/except isolation keeps a broken selector from killing
a run — but it also means a silently-empty field looks a lot like an attorney who genuinely
has no LinkedIn. A small assertion suite per firm (“this page should yield a non-empty law
school”) would have turned silent degradation into a loud failure.
Orchestration. Scheduled scraper runs are exactly the workload Airflow exists for — retries, per-firm task isolation, backfills, and a dependency graph that runs normalization after extraction rather than inside it. That’s the first thing I’d add.
What I took from it
The transformation is never the hard part. Ninety percent of the work in this project was upstream of any transformation: deciding what a valid value is, deciding what to do with something that isn’t one, and building the thing so that the answer is written down once instead of forty times.