Member of Technical Staff
Join a small, senior team, building the full on-device stack to achieve realtime local intelligence
ApplyThe role:
We’re hiring a Member of Technical Staff to work on various problems across our stack: lalamo and uzu.
This is a hybrid research and engineering role, where we expect the candidate to choose their own problems, come up with new ideas, conduct experiments and be up to take work that a moment requires - be it an engineering challenge or a fundamental research problem.
Most of the work will be released as open-source. The work that does not provide significant competitive advantage will be published as academic papers; see, for example, our speculative decoding approach.
Explore all rolesWhat you’ll do:
The current high-level tasks include post-training state-of-the-art LLMs in 10-50B range and improving their throughput on Apple devices using speculative decoding, quantization and other tricks. People switch between projects and everyone works on pretty much everything.
Representative problems include:
- Improving the tool-call accuracy in our harness using RLVR.
- Posttraining an LLM on several B200 nodes to control its reasoning effort better.
- Computing the decoding speed degradation bounds from KV cache quantization when using DFlash for speculative decoding.
- Constructing a sequential testing framework to improve the time-to-conclusion for the token throughput benchmarks.
What we’re looking for in a candidate:
- Solid knowledge of fundamentals of Machine Learning, Linear Algebra and Statistics
- Ability to write maintainable Python code
- High level of independence
- High level understanding of GPU architecture and programming model
Nice to have:
- Systems engineering experience, especially GPU/NPU programming or low-level programming experience
- Previous experience working with LLM inference or ML research
- Previous experience in PL theory, optimization theory or Metal.
Why join?
You’ll work on applied research that directly impacts how AI systems operate in real-world environments, not just benchmarks.
We are a fast-paced horizontal team with a lot of autonomy and trust.
We value technical depth and fast iteration. Competitive compensation + meaningful equity.
Why us?
Founded by proven entrepreneurs who built and scaled consumer AI leaders like Reface (300M users, a pioneer in Generative photo/video AI) and Prisma (100M MAU, a pioneer in on device AI photo enhancement).
Our team is small (16 people), senior, and deeply technical. We ship fast and own problems end-to-end. We’re advised by a former Apple Distinguished Engineer who worked on MLX, and backed by leading AI-focused funds and individuals.
Interested?
Join a small, senior team, building the full on-device stack to achieve realtime local intelligence
Apply