Post by Tobias Klein

I build client-owned patent data infrastructure: on-prem pipelines that turn USPTO/EPO PDF releases into byte-verifiable, deduplicated, ClickHouse-ready corpora for legal, search, analytics, and AI workflows.

The pipeline is designed to process any number of patents, from single PDF files to entire EPO/USPTO bulk data downloads. What that means is that it is built to handle 10,000+ files per run. But scale only matters if the same level of detail is preserved for every single patent. So let's look at one specific USPTO patent that Samsung Electronics Suwon-si (KR) Ltd. filed on April 10, 2024. It names Sangun Oh and Jaejeong Kim as inventors, was granted on March 24, 2026, and carries a 15-year term. Here is the actual output from the database for this USPTO patent. All text below is taken directly from the database page_text: Filed: Apr. 10, 2024 Applicant: Samsung Electronics Suwon-si (KR) Ltd. Inventors: Sangun Oh, Suwon-si (KR); Jaejeong Kim, Suwon-si (KR) Date of Patent: Mar. 24, 2026 Term: 15 Years Field of Classification Search: USPC Dl4/485 495 CPC GO6F 3/0481 GO6F 3/04845; GO6F 3/04817; GO6F 17/212; GO6F 9/44584; GOSD 2105/70; GOSD 2105/345; GIIB 23/40; HO4N 1/00183; HO4N 1/00193; HO4N 1/00251; HO4N 1/00257; HO4N 1/00172; HOAN 1/0018; HOAN 13/204 See application file for complete search history. References Cited U.S. PATENT DOCUMENTS: D696,264 12/2013 d Amore D14/485 D747,348 1/2016 Park D14/492 D772,249 11/2016 Choi D14/485 D804,521 12/2017 Deets Jr. D14/488 (Continued) And that is the point: a corpus of 10,000+ patents is still made up of 10,000 individual patents, each with its own filing dates, applicants, inventors, classifications, references, claims, page text, and exact source document. The scale does not replace the detail. The scale preserves it.

Post content