Job Summary
Own the on-device AI execution stack for ARM-based SoCs with dedicated neural accelerators — model conversion and quantization, runtime and delegate integration, and deployment of vision, LLM and VLM workloads within edge power and memory budgets.
Key Responsibilities
- Deploy and optimize neural network models on NPU, GPU and CPU backends using embedded inference runtimes (LiteRT/TensorFlow Lite, ONNX Runtime, ExecuTorch or vendor SDKs).
- Own the model conversion pipeline — graph capture, operator mapping, quantization (PTQ and QAT support), calibration and accuracy validation against reference.
- Diagnose and resolve unsupported operators, graph partitioning and fallback behaviour; work with vendor toolchains on accelerator limitations.
- Enable on-device LLM and VLM workloads — weight quantization, KV-cache management, prefill/decode optimization, memory footprint and token-throughput tuning.
- Benchmark and optimize inference latency, throughput, memory bandwidth and energy per inference; publish reproducible performance data.
- Integrate inference into real-time media and robotics pipelines with zero-copy tensor/buffer sharing.
- Build model deployment and regression tooling so accuracy and performance are tracked across releases.
- Advise product and data-science teams on model architecture choices suitable for edge accelerators.
Skill Requirements
- 8–10 years in software engineering, with 3+ years in on-device / edge AI deployment.
- Strong C++ and Python.
- Hands-on with embedded inference runtimes — LiteRT/TFLite, ONNX Runtime, ExecuTorch, TVM or vendor NPU SDKs — including delegate/execution-provider integration.
- Practical quantization expertise — INT8/INT4, per-channel schemes, calibration, accuracy recovery.
- Model formats and conversion tooling across PyTorch/TensorFlow to deployable graphs.
- Profiling on heterogeneous SoCs and reasoning about memory bandwidth as the dominant constraint.
- Understanding of CNN, transformer and modern vision/language model architectures.
Other Requirements
- On-device LLM/VLM deployment — llama.cpp class runtimes, speculative decoding, paged KV cache.
- Custom operator or kernel development for DSP/NPU/GPU (OpenCL, Vulkan compute, DSP intrinsics).
- Compiler-level work — MLIR, TVM, graph-level optimization passes.
- Vision pipeline integration and camera-to-inference zero-copy paths.
- MLOps for edge — model versioning, A/B evaluation, field accuracy monitoring.