Section A · Orient · Read first

Start Here

Interview prep for the Senior Data Scientist, Business Analytics role at a privacy-first AI company — where the defining challenge is building a data function without collecting the data most companies live on.

The role in one paragraph

The archetype this guide prepares you for is the first serious analytics hire at a privacy-first, uncensored consumer-AI company: a Senior Data Scientist, Business Analytics, reporting into business operations, fully remote, 7+ years of experience. The mandate is to stand up end-to-end data infrastructure, consolidate financial, product, and growth metrics into one place, support experiment tracking for product teams, and answer strategic business questions — all while being transparent about uncertainty and advocating against unnecessary data collection to protect the company's privacy commitments. This is not a modeling role. It is a build-the-foundation, own-the-truth, evangelize-the-culture role, at a company whose entire identity is not knowing what its users type.

Company-agnostic by design

This guide is deliberately not tied to one employer. It teaches the transferable spine of the role — privacy-first analytics, the data stack, the metric grammar of a freemium-plus-API-plus-token business, and the interview craft — so you can walk it into any company of this shape. If you're targeting a specific employer, do your own company research alongside it: know their product, their pricing, and where their money comes from cold.

The single most important reframe

Read this twice

Every other AI company mines the one asset this company deliberately throws away: the content of user conversations. If you walk into this interview treating the no-logging promise as an obstacle to work around, you will lose. The winning candidate treats it as the product, and proves you don't need user content to run a great data function. The line to internalize: "We're not flying blind — we're flying on instruments, deliberately privacy-safe ones."

Tactically, that reframe means: you lead with what you can measure (billing events, aggregate usage metadata, content-free product events, fully-public on-chain token data), you name the privacy-preserving techniques that make it rigorous (aggregation, k-anonymity, differential privacy, consented linkage), and you're honest about the real trade-offs it forces (no user-level funnels, no prompt corpus for evals). Chapter 02 is the whole answer. If you read nothing else, read that.

What the loop will test

This is calibrated from the posting, this class of company's culture, and how first-data-hire loops at profitable startups typically run. Expect some combination of:

  1. Company & motivation — do you actually understand the business, and do you believe in the privacy-first thesis? A crypto-libertarian founder screens hard for genuine alignment, not just competence. (This is where your own company research pays off.)
  2. The core design question — "how would you set up analytics here given we don't log prompts?" This is almost certainly asked. (Chapter 02.)
  3. Data stack / infrastructure — how you'd consolidate sources and choose tools as the first hire. (Chapter 03.)
  4. Metrics & business sense — freemium conversion, churn, NRR, token/staking metrics, inference unit economics. (Chapter 04.)
  5. SQL / technical screen — timed SQL on subscription and usage-shaped tables, plus evidence you write production-quality code. (Chapter 07.)
  6. Experimentation — designing and tracking experiments without user-level identity. (Chapter 05.)
  7. Communication & behavioral — explaining uncertainty to executives, telling first-hire / ambiguity stories, and defending "advocate against data collection" with a straight face. (Chapter 08.)

The guide in reading order

Four sections. The numbering is the order to read in.

Section A — Orient

ChapterWhy
01 · The Role DecodedThe posting mapped line-by-line to what it means, the first-data-hire archetype, comp, and who you report to

Section B — The Core Challenge

ChapterWhy
02 · Analytics Without SurveillanceThe defining question. What you can/can't measure and the privacy-preserving toolkit
03 · Building the Data StackWarehouse, ingestion, dbt, BI — the "establish end-to-end infrastructure" mandate, with a 30-60-90
04 · Metrics That MatterFreemium + API + token metrics, and the inference-COGS margin story
05 · ExperimentationExperiment tracking and privacy-safe A/B for a fast-iterating product
06 · On-Chain AnalyticsThe public token data most candidates miss, and the consented-linkage rule

Section C — Technical & Communication

ChapterWhy
07 · SQL & Technical ScreenSQL on subscription- and usage-shaped tables, plus production-code expectations
08 · Communicating UncertaintyRecommendations with confidence levels, and defending the privacy advocacy clause

Section D — Execution

ChapterWhy
09 · Practice Questions~30 calibrated Q&A with drill mode
10 · Day-Of & 30-60-90The plan to walk in with, questions to ask, closing. Reread morning of

Study schedule

Three calibrations — the same paths are on the hub:

  • 4+ days: one section per day, and actually use a product of this kind — sign up, run a few prompts, read its privacy policy, open a token dashboard on Dune. Nothing lands like first-hand product experience in the room.
  • 2 days: 00, 01, 02, 03, 04, 07 (drill), 09 (drill), 10. Skim 05, 06, 08.
  • An evening: 00, 01, all of 02, the 30-60-90 in 03 and 10, drill 09.

A note on honesty

When you research your target company, expect several facts to be moving targets or self-reported — funding, ARR, user counts, token prices, and leadership names all drift, and startups quote generously. A candidate who says "their reported ARR is around that range, though that's self-reported" sounds more credible than one who states it as gospel. That instinct — being precise about what you know versus what you're inferring — is literally the job ("be transparent about analytical uncertainty"). Model it from the first conversation.

Verify before you walk in

The morning of, spend ten minutes refreshing your target company's current pricing, its live model list, any token price and market cap, and the exact job posting again in case it changed. These are cheap points and stale facts are unforced errors.

What winning looks like

You don't need to be the strongest pure modeler in the loop. You need to be the candidate who:

  1. Clearly understands what the company is, how it makes money, and why the privacy stance is a strategic choice — not a marketing gimmick.
  2. Can answer "how do you build analytics here without logging users" with a structured, specific plan — safe-vs-unsafe data, a real stack, and the privacy techniques that make it rigorous.
  3. Writes clean SQL fast on subscription and usage-shaped tables, without flinching at window functions.
  4. Talks about metrics like an operator — free→paid, NRR, the smiling curve, staking as prepaid demand, inference as COGS — not like a textbook.
  5. Knows the token economy exists as a public data source, and can describe analyzing it without ever de-anonymizing a user.
  6. Gives recommendations with a confidence level and a "what would change my mind" sentence — and genuinely believes in advocating against unnecessary data collection.

If you can do those six things on demand, you're in. Turn to 01 · The Role Decoded.