I maintain a few datasets to help other researchers and save their efforts and they are mostly used in my papers. I also maintain several software that help people to be more productive. My most starred Github repo helped thousands of people in both academia and industry
Comprehensive, high-precision U.S. patent-to-scholarly paper linkage and dual-disclosure dataset (61.7M links, 1.29M patents, 3.74M OpenAlex papers). Extends Marx & Fuegi beyond 2021 through 2026 with 97.7% twin precision.
ratex
Ultra-fast, self-contained pure-Rust TeX engine and toolchain. Compiles documents in milliseconds with an embedded LaTeX format and 24,000+ packages.
CIK to CUSIP Mapping
Provide linking files between CIK and CUSIP using 13G and 13F filings.
USPTO full text database
Provide OCR full text data for pre-1975 USPTO patents. They offer great improvements in quality and coverage than those in Google Patents
Name Matching
Algorithm to match firm names based on string similarities
Replace and Delete (rd)
Extremely fast command line utility to replace and delete strings in text files
Fuzzy Process (fuzzprocess)
Deep-learning approach to find nearest K matches for two sets of names