Multimodal AI Model Optimization Research Engineer at Tavus — AI Jobs | Neura Market
    Neura Market
    Neura Market
    /Jobs
    Marketplace
    Directories
    Resources
    AI JobsMultimodal AI Model Optimization Research Engineer
    Tavus

    Multimodal AI Model Optimization Research Engineer

    Tavus

    Remote

    Marketplace

    • Prompts
    • Workflows
    • Agent Hub
    • Workflow Packs
    • Categories
    • Marketplace

    Directories

    • AI Tools Directory
    • ChatGPT
    • Claude
    • Gemini
    • Cursor
    • Grok
    • DeepSeek
    • Perplexity
    • CoPilot
    • Midjourney
    • Stable Diffusion
    • MCP Servers
    • .md Directory
    • All Directories

    Free Tools

    • AI Text Humanizer
    • AI Content Detector
    • Workflow Generator
    • Model Comparison
    • AI Pricing Calculator
    • AI Benchmarks
    • ROI Calculator
    • All Free Tools

    Resources

    • AI News
    • Blog
    • AI Answers
    • Error Solutions
    • AI Tutorials
    • AI Agent Guides
    • AI Models
    • AI Research Papers
    • Integrations
    • Alternatives
    • n8n vs Zapier
    • Make vs Zapier
    • n8n vs Make
    • Resource Library
    • Documentation
    • API Access to Our Data

    Community

    • AI Newsletter
    • AI Jobs
    • AI Events
    • AI Companies
    • Start Selling
    • Sell n8n Workflows
    • Sell AI Agents
    • Sell Prompts
    • Creator Guide
    • Advertise
    • Affiliates

    Company

    • About
    • Contact
    • Help
    • Careers
    • Pricing
    • Terms
    • Privacy
    • License
    • DMCA

    The #1 Newsletter in AI

    Weekly updates, news, and content that matter.

    Neura Market Logoneuramarket

    © 2026 Neura Market. All rights reserved.

    Full-time
    Remote
    6/15/2026
    Apply

    About This Role

    About Us

    At Tavus, we're building the human layer of AI. Our mission is to make human-AI interaction as natural as face-to-face interaction, enabling the human touch where it has been previously unscalable.

    We achieve this through pioneering research in multimodal AI for modeling human-to-human communication (language, audio, and video), as well as generating audio-visual avatar behavior. Our models power everything from text-to-video AI avatars to real-time conversational video experiences across industries like healthcare, recruiting, sales, and education.

    By enabling AI to see, hear, and communicate with human-like authenticity, we're creating the foundation for the next generation of AI employees, assistants, and companions.

    We are a Series B company backed by top investors, including Sequoia, Y Combinator, and Scale VC. Join us in driving the future of human-AI interaction.

    The Role

    We’re looking for an experienced Research Scientist/Engineer with a focus on model optimization to join our core AI team.

    Our ideal partner-in-crime thrives in startup environments, is comfortable prioritizing independently, and is willing to take calculated risks. We’re moving fast and looking for people who can help pave the path.

    Your Mission

    • Take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization

    • Own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality

    • Partner closely with researchers and engineers to turn new ideas into deployable systems

    Requirements

    • Strong experience in deep learning using PyTorch

    • Hands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision

    • Understanding of efficient architectures such as low-rank adapters

    • Strong understanding of inference performance and GPU/accelerator fundamentals

    • Strong Python coding skills and reliable research engineering practices

    • Experience working with large models and datasets in cloud environments

    • Ability to read ML papers, reproduce results, and adapt ideas

    • Clear communication and collaboration skills

    Preferred Experience

    • Optimization of diffusion models, video/audio generative models, or large language models

    • Experience with real-time or streaming systems (low-latency APIs, WebRTC, streaming TTS/video)

    • Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA

    • Experience writing custom Triton/CUDA kernels or low-level performance tuning

    • Experience with experiment tracking, benchmarking, and profiling at scale

    • Prior experience in research engineering or applied science roles

    Location

    This position is preferably hybrid in San Francisco, with relocation support offered. Remote candidates are also considered.

    Benefits

    When you join Tavus, you’re joining a family. We offer flexible work schedules, unlimited PTO, competitive healthcare and gear stipends, and a collaborative environment focused on learning and impact.

    Culture & Diversity

    We are not looking for cultural fits — we are looking for culture creators. Diversity drives our success, and we combine varied backgrounds, skills, and perspectives to build the best experiences for our clients..

    Tasks

    • •Take cutting-edge research models and make them fast, efficient, and production-ready using sparsification, distillation, and quantization
    • •Own the optimization lifecycle for key models: define metrics, run experiments, and benchmark trade-offs across latency, cost, and quality
    • •Partner closely with researchers and engineers to turn new ideas into deployable systems
    • •Strong experience in deep learning using PyTorch
    • •Hands-on experience with model optimization and compression, including knowledge distillation, pruning/sparsification, quantization, and mixed precision
    • •Understanding of efficient architectures such as low-rank adapters
    • •Strong understanding of inference performance and GPU/accelerator fundamentals
    • •Strong Python coding skills and reliable research engineering practices
    • •Experience working with large models and datasets in cloud environments
    • •Ability to read ML papers, reproduce results, and adapt ideas
    • •Clear communication and collaboration skills
    • •Optimization of diffusion models, video/audio generative models, or large language models
    • •Experience with real-time or streaming systems (low-latency APIs, WebRTC, streaming TTS/video)
    • •Familiarity with TensorRT, ONNX Runtime, TVM, Triton, or XLA
    • •Experience writing custom Triton/CUDA kernels or low-level performance tuning

    Perks & Benefits

    Health insuranceUnlimited vacationPaid time offRelocation assistance

    Skills & Tech Stack

    PythonPyTorchDiffusion ModelsCUDATriton

    Roles

    Research EngineerEngineer

    Topics

    Engineering, Product, & Design

    Related AI Jobs

    Databricks

    Staff Backend Software Engineer- (AI Platform)

    Databricks·Full-time·San Francisco, California
    Engineering
    Snowflake

    Staff Security Engineer - Threat Detection

    Snowflake·Full-time·US, Remote
    Engineering
    Snowflake

    Software Engineer - Openflow

    Snowflake·Full-time·US-CA-Menlo Park
    Engineering
    Decagon

    Product Manager, Enterprise Agent Platform

    Decagon·Full-time·San Francisco

    $232,000 - $290,000/yr

    Product
    Decagon

    IT Engineer

    Decagon·Full-time·New York City

    $96,000 - $132,000/yr

    Engineering
    Weaviate

    Director of Product

    Weaviate·Full-time·CET, GMT or EST timezones
    Product
    ← Back to all jobs