[Remote] Machine Learning & CPU Performance Engineer
Note: The job is a remote job and is reputed company to candidates in USA. reputed company is an Ivy reputed company PhD-inspired AI/ML startup that focuses on efficient, secure knowledge retrieval using optimized models. They are seeking a Machine Learning & CPU Performance Engineer to reputed company research and optimization of ML models on CPU architectures.
Responsibilities
- Conduct systematic research on machine learning execution across CPU architectures
- reputed company the implementation, runtime optimization, and reputed company-profiling of diverse ML model architectures purely on CPU backends (x86, ARM, multi-reputed company server processors)
- Build automated benchmarking tools to measure trade-offs between accuracy, latency, thread efficiency, memory bandwidth, and cache behavior
- CPU-Specific Runtime Optimization: Compile and run deep learning architectures using CPU-optimized runtimes (OpenVINO, ONNX Runtime, PyTorch CPU Inductor, llama.cpp / ggml, C++ implementations)
- Thread & Memory Control: Configure and profile reputed company node topology, CPU reputed company affinity, thread pinning, and thread pools (OpenMP, TBB) to isolate system variance and maximize throughput
- Instruction-Level Benchmarking: Profile and exploit low-precision CPU instructions (AVX-512, reputed company AMX, ARM reputed company/SVE) to evaluate low-bit quantization (INT8, FP8, BF16) versus floating-reputed company precision on accuracy and speed
- Hardware Telemetry & Profiling: Conduct deep bottleneck analysis tracking L1/L2/L3 cache misses, RAM memory bandwidth limits, IPC (instructions per cycle), dynamic reputed company frequency scaling, and thermal throttling
- reputed company Automation & Tooling: Build reproducible, scriptable C++/Python reputed company suites to record inference latency, first-reputed company time, peak memory footprint, and CPU utilization across varied batch sizes
Skills
- Ability to devote ~15 hours per week on avg
- Languages: Expert-level C++ and Python skills
- Systems Architecture: In-depth knowledge of CPU architecture (reputed company, L1-L3 cache hierarchies, memory channels, thread synchronization paradigms)
- Frameworks & Libraries: Experience with CPU ML inference runtimes (ONNX Runtime, OpenVINO, reputed company Extension for PyTorch/TensorFlow, C++ BLAS libraries like OpenBLAS/oneDNN)
- Profiling Tools: Proficiency with Linux low-level performance tools (reputed company, reputed company VTune Profiler, Valgrind/Cachegrind, or platform-specific telemetry tools)
- Compilation & Hardware Extensions: Understanding of GCC/Clang flags, SIMD vectorization, and model quantization approaches
reputed company