Data Engineer / Data Scientist — Brand Index
Optimly · Remote (Seattle, WA preferred) · Full-time · $130k–$200k + 0.3%–1.0% equity
About Optimly
Optimly is where AI models go to learn about brands and where brands go to learn about AI models.
Half of the product is an app where a brand sees how frontier models describe it. The other half is the Brand Index: a public, machine-readable catalogue of tens of thousands of brands, built so that AI crawlers and agent backends can read it directly.
The index is our distribution engine. It's also our messiest data problem.
The role
You'd own the index end to end: where the data comes from, how it's cleaned and resolved, how it's stored, and how it's served to both humans and machines. This includes:
- Ingestion and extraction
- Identity and data quality
- Machine readability
- Growth
What we're looking for
- Strong Python and SQL.
- Real experience with messy web-scale data: scraping, extraction, deduplication, entity resolution, canonical identity.
- Statistical judgment.
- Working knowledge of how AI retrieval actually works or a demonstrated ability to learn it fast.
Nice to have: SEO or search-quality background, public data or directory products, evaluation of LLM extraction quality, large-scale crawler operations.
How we work
We use AI heavily and expect you to, as well. Day-to-day work happens in agentic coding tools. For this role that cuts both ways: you'll use models to build pipelines, and you'll also be the person who defines how we know a model-driven pipeline is right. We want someone who treats an LLM as a component with a measured error rate — not as an oracle, and not as something to avoid.
We're small. You'll see the whole system and talk to customers.
To apply
Send us a short note about the dirtiest dataset you've had to make trustworthy.