Skip to main content
Remote / SF / EuropeFull Time

Member of Technical Staff

Join a small, senior team, building the full on-device stack to achieve realtime local intelligence

Apply

The role:

We’re hiring a Member of Technical Staff to work on various problems across our stack: lalamo and uzu.

This is a hybrid research and engineering role, where we expect the candidate to choose their own problems, come up with new ideas, conduct experiments and be up to take work that a moment requires - be it an engineering challenge or a fundamental research problem.

Most of the work will be released as open-source. The work that does not provide significant competitive advantage will be published as academic papers; see, for example, our speculative decoding approach.

Explore all roles

What you’ll do:

The current high-level tasks include post-training state-of-the-art LLMs in 10-50B range and improving their throughput on Apple devices using speculative decoding, quantization and other tricks. People switch between projects and everyone works on pretty much everything.

Representative problems include:

  • Improving the tool-call accuracy in our harness using RLVR.
  • Posttraining an LLM on several B200 nodes to control its reasoning effort better.
  • Computing the decoding speed degradation bounds from KV cache quantization when using DFlash for speculative decoding.
  • Constructing a sequential testing framework to improve the time-to-conclusion for the token throughput benchmarks.

What we’re looking for in a candidate:

  • Solid knowledge of fundamentals of Machine Learning, Linear Algebra and Statistics
  • Ability to write maintainable Python code
  • High level of independence
  • High level understanding of GPU architecture and programming model

Nice to have:

  • Systems engineering experience, especially GPU/NPU programming or low-level programming experience
  • Previous experience working with LLM inference or ML research
  • Previous experience in PL theory, optimization theory or Metal.

Why join?

You’ll work on applied research that directly impacts how AI systems operate in real-world environments, not just benchmarks.

We are a fast-paced horizontal team with a lot of autonomy and trust.

We value technical depth and fast iteration. Competitive compensation + meaningful equity.

Why us?

Founded by proven entrepreneurs who built and scaled consumer AI leaders like Reface (300M users, a pioneer in Generative photo/video AI) and Prisma (100M MAU, a pioneer in on device AI photo enhancement).

Our team is small (16 people), senior, and deeply technical. We ship fast and own problems end-to-end. We’re advised by a former Apple Distinguished Engineer who worked on MLX, and backed by leading AI-focused funds and individuals.

Interested?

Join a small, senior team, building the full on-device stack to achieve realtime local intelligence

Apply