Jobiglo

Sin resultados.

AI Research Engineer (Kernel & Inference Optimization)

Tether.io · Andorre

Nuevo Remote
Remote 🇬🇧 English
Metal Shading Language (MSL) Model serving frameworks Tensor Parallelism Pipeline Parallelism Expert Parallelism Pruning Flash attention KV cache Speculative decoding Diffusion models Vision transformers

Descripcion del puesto

About the role

As a member of Tether's AI model team, you will lead the design and optimisation of model serving and inference pipelines for advanced AI systems. The focus is on delivering high‑throughput, low‑latency, and memory‑efficient performance across a range of devices, from edge platforms to large GPU clusters.

Key responsibilities

  • Design and deploy state‑of‑the‑art model serving architectures that maximise throughput while minimising latency and memory usage.
  • Build, run and monitor controlled inference tests in simulated and live production environments, tracking latency, throughput, memory consumption and error rates.
  • Identify bottlenecks in the serving pipeline and implement kernel‑level and system‑level optimisations for resource‑constrained devices.
  • Collaborate with cross‑functional teams to integrate optimised inference frameworks into edge and on‑device production pipelines.
  • Prepare high‑quality test datasets and simulation scenarios that reflect real‑world deployment challenges on low‑resource hardware.

Required profile

  • Degree in Computer Science or related field; PhD in NLP, Machine Learning or a related discipline is preferred.
  • Proven track record in AI R&D with publications in top conferences.
  • Deep knowledge of Metal Shading Language (MSL) and experience writing custom compute shaders.
  • Extensive experience in low‑level kernel optimisation and inference optimisation on mobile devices.
  • Strong understanding of modern model serving architectures and distributed inference techniques.

Required skills

  • Metal Shading Language (MSL) and custom GPU kernel development for mobile devices.
  • Low‑level kernel optimisation and inference optimisation on resource‑constrained hardware.
  • Model serving frameworks, Tensor Parallelism, Pipeline Parallelism, Expert Parallelism.
  • Techniques such as pruning, quantisation, flash attention, KV cache and speculative decoding.
  • Knowledge of diffusion models and vision transformers.

Questions fréquentes

Le salaire n'est pas communiqué publiquement par le recruteur. Vous pouvez postuler et négocier directement avec Tether.io.
Cliquez sur "Postuler maintenant" en haut de la page. Vous pouvez importer votre CV en 1 clic — Jobiglo extrait automatiquement vos informations et postule pour vous.

Por que reporta esta oferta?

Gracias por su reporte. Revisaremos esta oferta.

Postula en 30 segundos

Ingresa tu email para postular. Se creara una cuenta automaticamente.

Al continuar, aceptas nuestras condiciones de uso.

Ya tienes cuenta? Iniciar sesion

💬 Escríbenos en Telegram Chatear por WhatsApp

Publicado hace 9 horas

Expira en 1 mes

3 vistas · 0 interested

Aumenta tus posibilidades

Sube tu CV: te propondremos las ofertas que coinciden con tu perfil.

Analizando tu CV...

Tether.io

Andorre