Jobiglo

No results.

AI Research Engineer (Kernel & Inference Optimization)

Tether.io · Andorre

New Remote
Remote 🇬🇧 English
Metal Shading Language (MSL) Model serving frameworks Tensor Parallelism Pipeline Parallelism Expert Parallelism Pruning Flash attention KV cache Speculative decoding Diffusion models Vision transformers

Job description

About the role

As a member of Tether's AI model team, you will lead the design and optimisation of model serving and inference pipelines for advanced AI systems. The focus is on delivering high‑throughput, low‑latency, and memory‑efficient performance across a range of devices, from edge platforms to large GPU clusters.

Key responsibilities

  • Design and deploy state‑of‑the‑art model serving architectures that maximise throughput while minimising latency and memory usage.
  • Build, run and monitor controlled inference tests in simulated and live production environments, tracking latency, throughput, memory consumption and error rates.
  • Identify bottlenecks in the serving pipeline and implement kernel‑level and system‑level optimisations for resource‑constrained devices.
  • Collaborate with cross‑functional teams to integrate optimised inference frameworks into edge and on‑device production pipelines.
  • Prepare high‑quality test datasets and simulation scenarios that reflect real‑world deployment challenges on low‑resource hardware.

Required profile

  • Degree in Computer Science or related field; PhD in NLP, Machine Learning or a related discipline is preferred.
  • Proven track record in AI R&D with publications in top conferences.
  • Deep knowledge of Metal Shading Language (MSL) and experience writing custom compute shaders.
  • Extensive experience in low‑level kernel optimisation and inference optimisation on mobile devices.
  • Strong understanding of modern model serving architectures and distributed inference techniques.

Required skills

  • Metal Shading Language (MSL) and custom GPU kernel development for mobile devices.
  • Low‑level kernel optimisation and inference optimisation on resource‑constrained hardware.
  • Model serving frameworks, Tensor Parallelism, Pipeline Parallelism, Expert Parallelism.
  • Techniques such as pruning, quantisation, flash attention, KV cache and speculative decoding.
  • Knowledge of diffusion models and vision transformers.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Tether.io.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Why are you reporting this job?

Thank you for your report. We will review this job.

Apply in 30 seconds

Enter your email to apply. An account will be created automatically.

By continuing, you accept our terms of use.

Already have an account? Login

💬 Chat with us on Telegram Chat on WhatsApp

Published 6 hours ago

Expires 1 month from now

1 views · 0 interested

Boost your chances

Upload your CV — we will match you with relevant openings.

Analyzing your CV...

Tether.io

Andorre