ML Engineer, Inference & Optimization at Pika — AI Jobs | Neura Market
    Neura Market
    Neura Market
    /Jobs
    Marketplace
    Directories
    Resources
    AI JobsML Engineer, Inference & Optimization
    Pika

    ML Engineer, Inference & Optimization

    Pika

    Palo Alto HQ

    Marketplace

    • Prompts
    • Workflows
    • Agent Hub
    • Workflow Packs
    • Categories
    • Marketplace

    Directories

    • AI Tools Directory
    • ChatGPT
    • Claude
    • Gemini
    • Cursor
    • Grok
    • DeepSeek
    • Perplexity
    • CoPilot
    • Midjourney
    • Stable Diffusion
    • MCP Servers
    • .md Directory
    • All Directories

    Free Tools

    • AI Text Humanizer
    • AI Content Detector
    • Workflow Generator
    • Model Comparison
    • AI Pricing Calculator
    • AI Benchmarks
    • ROI Calculator
    • All Free Tools

    Resources

    • AI News
    • Blog
    • AI Answers
    • Error Solutions
    • AI Tutorials
    • AI Agent Guides
    • AI Models
    • AI Research Papers
    • Integrations
    • Alternatives
    • n8n vs Zapier
    • Make vs Zapier
    • n8n vs Make
    • Resource Library
    • Documentation
    • API Access to Our Data

    Community

    • AI Newsletter
    • AI Jobs
    • AI Events
    • AI Companies
    • Start Selling
    • Sell n8n Workflows
    • Sell AI Agents
    • Sell Prompts
    • Creator Guide
    • Advertise
    • Affiliates

    Company

    • About
    • Contact
    • Help
    • Careers
    • Pricing
    • Terms
    • Privacy
    • License
    • DMCA

    The #1 Newsletter in AI

    Weekly updates, news, and content that matter.

    Neura Market Logoneuramarket

    © 2026 Neura Market. All rights reserved.

    Senior-level / Expert
    Full-time
    On-site
    6/23/2026
    Apply

    About This Role

    About the Role

    We are seeking Senior/Staff level Inference Engineers to accelerate the performance of Pika's AI-driven products. In this highly technical role, you will operate at the intersection of cutting-edge inference acceleration, GPU parallelism, advanced model deployment, and video generation technologies. Your expertise will drive significant improvements to model speed and efficiency, ensuring our creative AI systems deliver industry-leading user experiences at scale.

     

    You will design and optimize inference pipelines, implement state-of-the-art acceleration techniques, and work closely with researchers and engineers across the team to push the boundaries of what’s possible in real-time AI deployment. Your efforts will play a foundational role in powering the next generation of Pika’s video and language models.

     

    What You’ll Do

    • Accelerate Inference: Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.

    • Maximize GPU Parallelism: Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.

    • Programming for Performance: Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.

    • Advance AI Deployment: Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.

    • Improve Training Efficiency: (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.

    • Technical Excellence: Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.

     

    What We’re Looking For

    • Experience: 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale.

    • Inference Mastery: Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.

    • GPU & Parallelism: Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.

    • AI Domain Knowledge: Familiarity with video generation (videogen) models and large language models (LLMs).

    • Collaboration: Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.

    • Ownership Mindset: Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.

    • Bonus: Experience in enhancing training efficiency, stability, or resource optimization for large models.

     

    Nice to Have

    • Experience with high-throughput video or real-time streaming model deployment

    • Familiarity with distributed training and optimization toolkits

    • Contributions to open source projects in AI infrastructure or deep learning compilers

    • Startup or rapid prototyping experience

     

    What We Offer

    • Competitive salary in the AI industry

    • Equity in a fast-growing startup shaping the future of AI

    • Comprehensive health benefits, monthly stipends, company retreats

    • A supportive and collaborative office culture—we’re all building and launching together

     

    About Pika

    At Pika, we're crafting a future where video creation is seamless, intuitive, and universally accessible. Our mission is to empower creativity by breaking down technical barriers using the transformative power of AI. We’re a tight-knit, energetic team based in Palo Alto, CA, valuing efficiency, curiosity, and the ambition to make a meaningful impact on the world.

     

    We work from our Palo Alto office 3–5 days a week and welcome applicants who are eager to contribute onsite.

    Tasks

    • •Accelerate Inference : Lead and implement advanced inference acceleration techniques, including attention optimization and quantization for efficient model serving.
    • •Maximize GPU Parallelism : Engineer and optimize GPU strategies across tensor, sequence, and pipeline parallelism (TP, SP, PP) for maximal efficiency and scalability.
    • •Programming for Performance : Develop and optimize high-performance computing kernels and distributed workloads using CUDA and NCCL.
    • •Advance AI Deployment : Collaborate with research and engineering teams to bring state-of-the-art videogen and large language models into production.
    • •Improve Training Efficiency : (Bonus) Contribute to improvements in model training speed, stability, and resource utilization as part of our deployment lifecycle.
    • •Technical Excellence : Drive rigorous code reviews, participate in technical discussions, and mentor fellow engineers on best practices in inference and GPU programming.
    • •Experience : 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale.
    • •Inference Mastery : Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks.
    • •GPU & Parallelism : Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference.
    • •AI Domain Knowledge : Familiarity with video generation (videogen) models and large language models (LLMs).
    • •Collaboration : Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions.
    • •Ownership Mindset : Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment.
    • •Bonus : Experience in enhancing training efficiency, stability, or resource optimization for large models.

    Perks & Benefits

    Competitive salary in the AI industryEquity in a fast-growing startup shaping the future of AIComprehensive health benefits, monthly stipends, company retreatsA supportive and collaborative office culture—we’re all building and launching together

    Skills & Tech Stack

    CUDA

    Roles

    ML EngineerEngineer

    Topics

    Research

    Related AI Jobs

    OpenAI

    Technical Commodity Manager - Robotics

    OpenAI·Full-time·San Francisco
    Research
    Exa AI

    Research, Singapore

    Exa AI·Full-time·Singapore
    Research
    OpenAI

    Field Engineer

    OpenAI·Full-time·San Francisco
    Research
    Scale AI

    VP, Research

    Scale AI·Full-time·San Francisco, CA; New York, NY
    Research
    Together AI

    Research Engineer, Large-Scale Training

    Together AI·Full-time·San Francisco

    $200,000 - $290,000/yr

    Research
    Scale AI

    Manager, Research Scientist

    Scale AI·Full-time·San Francisco, CA; New York, NY
    Research
    ← Back to all jobs