Speak
Assessment Design Lead
Remote
Posted August 10, 2026
Job description
ABOUT US
Our mission is to reinvent the way people learn, starting with language.
Learning a language can change a life by opening doors to new cultures, careers, and communities. Two billion people around the world are actively trying to learn a language, but the best way to learn (one-on-one tutoring) is hard to access at scale and hasn’t been meaningfully improved in decades. Speak is building a human-level, AI-powered tutor in your pocket: a conversation-first experience that lets learners actually speak, get instant feedback, and progress through carefully designed lessons. The result is a complete path from beginner to confident speaker across multiple languages.
Speak first launched in South Korea in 2019, where Speak has now become the number one language learning app, and we now serve learners across many markets and 15+ languages. Speak is one of the world’s leading AI companies, with over $150m raised in venture investment from OpenAI, Accel, Founders Fund, Khosla Ventures, and more, with a distributed team across San Francisco, Seoul, Tokyo, Taipei, and Ljubljana.
ABOUT THIS ROLE
Speak cares deeply about learners actually learning and improving with Speak. We have a dedicated Proficiency team to own how we measure learning efficacy and speaking proficiency, from unit-level mastery checks to standalone proficiency tests to onboarding placement, all in service of understanding users’ proficiency levels and learning gains in an accurate, transparent and actionable manner. Because everything happens remotely and asynchronously in the app, keeping scores fair, stable over time, and resistant to gaming is both hard and genuinely interesting.
We're looking for an Assessment Design Lead: the person who defines what we measure, why, how, and signs off on whether our assessments actually measure the right thing. You will be staffed on the Proficiency team and report to the Head of Learning Design and Curriculum. If solving reliable, at-scale speaking assessment excites you and you want to directly shape Speak's efficacy story, we'd love to hear from you.
WHAT YOU'LL BE DOING
- Own Assessment Design — Define what Speak measures, why, and how often across three distinct assessment types (Curriculum Mastery Assessment, Proficiency Test, Placement Test) — with the Proficiency Test as the immediate focus, expanding to the other two as the pod's priorities evolve. Keep constructs (fluency, pronunciation, grammar, task achievement) clearly separated and each aligned to CEFR or a comparable speaking proficiency standard, so no single assessment conflates domains it wasn't designed to measure.
- Define Constructs & Build the Rubric/Blueprint Layer — Translate fuzzy goals like "measure fluency" or "measure pronunciation" into concrete, scoreable constructs, item blueprints, and rubrics that an item writer can generate items against and an ML Engineer can build a grading model against.
- Own Validity & the Quality Bar — Sign off on content validity for every assessment that ships. Decide what "mastery" or a passing score operationally means, catch cases where an assessment is measuring the wrong thing before it ships, own the rubric/rater guidelines behind the human-labeled data our ML scoring models are evaluated against, and audit items/rubrics for bias across learner subgroups. Making sure scores stay comparable as the assessment evolves and stay meaningful against attempts to game an unproctored test.
- Design and run the validity evidence plan — so validity is built into the process rather than checked only after launch. This includes concurrent/criterion studies benchmarking Speak’s assessments against external proficiency measures (CEFR-anchored exams, expert human ratings), so we can say what a Speak Score means in terms the outside world already trusts.
- Partner Tightly with Product and ML — Work closely with the Product Manager and ML Engineers on automated scoring, calibration, and feedback generation. You own the construct and quality bar, they own the model. Neither works without the other, and the loop between you is the product.
WHAT WE'RE LOOKING FOR
MUST-HAVES
- Assessment/Psychometric Design: 4+ years designing rubrics, blueprints, and item specs for a real, shipped language assessment product (or equivalent depth in closely related psychometric/measurement work) — not just academic theory. Can explain reliability and validity in plain language and knows how to catch a test that's measuring the wrong construct.
- Language Proficiency Domain Expertise: Deep familiarity with frameworks like CEFR (or ACTFL, IELTS/TOEFL band descriptors) and what separates "did you learn what we taught you" from "how good is your speaking overall."
- Fairness Across Learner Populations: Can identify whether an item or rubric unfairly penalizes specific L1 backgrounds or accents (differential item functioning) — essential for a speech-based test serving learners across dozens of native languages.
- Translates Qualitative → Technical: Can turn a construct like "pronunciation quality" into something concrete enough for an ML engineer to build a scoring pipeline against, without either oversimplifying or getting lost in academic nuance.
- Quantitative Rigor: Comfortable running or interpreting the statistics behind a rubric or rater system — inter-rater reliability (e.g., Cohen's/Fleiss' kappa), classical test theory, and basic IRT concepts — enough to know whether a scoring system is actually reliable, not just plausible.
- Ownership of Quality Bar: Comfortable being the sign-off authority on content validity — makes the call clearly and follows through on it, rather than deferring to data alone or product pressure to ship.
- AI Fluency & Judgment: Uses AI tools directly in their own workflow (e.g., drafting item variants, testing rubric language, exploring construct definitions) and has real judgment about when AI-generated output is precise enough to ship vs. needs a human rewrite — distinct from spec'ing work for the ML Engineer to build.
- Comfort with Ambiguity: Comfortable operating in a 0-to-1 environment. Can wear multiple hats, take a fuzzy goal and turn it into a concrete plan, communicate tradeoffs clearly, and keep momentum without waiting for perfect clarity or team setup
NICE-TO-HAVES
- Speech/pronunciation science background — can own pronunciation frameworks and L2-specific error taxonomy directly
- Familiarity with adaptive testing or IRT-adjacent concepts (even if not the primary psychometrician)
- Experience at a large-scale language testing organization or similar high-rigor assessment environment
- Experience thriving in an EdTech startup environment, especially in a newly forming team or 0-to-1 mandate
- Advanced degree (Master's or PhD) in psychometrics, measurement, applied linguistics, SLA, or a related quantitative field; track record of shipped assessment work still matters more
- Has authored technical/validity reports or published assessment research
HOW WE WORK
This role is designed to be highly collaborative with our Assessment ML Engineer. Success depends on a tight loop where constructs, rubrics, and model outputs co-evolve together — from the earliest fuzzy construct through to a shipped scoring model — rather than a one-time handoff from a finished spec to an implementation.
WHY WORK AT SPEAK
1. Join a fantastic, tight-knit team at the right time: we're growing very quickly, we've most recently raised our Series C from some of the top investors in the valley, and we've achieved product-market fit in our initial markets. You'd join at a magical time when a single person could significantly change the course of the company.
2. Do your life's work with people you’ll love working with: we care strongly about our craft and want every person at Speak to feel like they're growing every day. We believe in the idea that working with people you both enjoy and have respect for makes everything better. We hire thoughtfully and only work with people we admire deeply.
3. Global in nature: We're live in over 40 countries and launching in a number of new markets soon. We have dedicated offices in San Francisco, Ljubljana, Seoul, Tokyo, and Taipei, and you’ll have the opportunity to talk to users in each of these regions on a regular basis as well as travel.
4. Impact people's lives in a major way: Learning a language is one of the single most life-changing skills one can learn, and right now 99% of people never achieve their goal because the process is broken. We’re helping millions of people achieve their goals and improve their lives.
Speak does not discriminate based upon race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics.