Advanced AI Inference Software Developer
## Full Description Company Overview AGI is a leading innovator in multi-modal inference solutions, revolutionizing how AI systems perceive and interact with the world. Our team pushes the limits of inference performance to provide exceptional user experiences across various applications and devices. Job Description We are seeking talented Inference Engineers to develop high-performance inference software for diverse neural models, typically in C/C++. You will design, prototype, and evaluate new inference engines and optimization techniques, participate in deep-dive analysis and profiling of production code, optimize inference performance across various platforms, and collaborate closely with research scientists to bring next-generation neural models to life. Key Responsibilities • Develop high-performance inference software for a wide range of neural models • Design, prototype, and evaluate new inference engines and optimization techniques • Participate in deep-dive analysis and profiling of production code • Optimize inference performance across various platforms • Collaborate closely with research scientists to bring next-generation neural models to life BASIC QUALIFICATIONS • 3+ years of non-internship professional software development experience • 2+ years of non-internship design or architecture experience of new and existing systems • Experience programming with at least one software programming language • Bachelor's degree in Computer Science, Computer Engineering, or related field • Strong C/C++ programming skills • Solid understanding of deep learning architectures PREFERRED QUALIFICATIONS • 3+ years of full software development life cycle experience • Experience with inference frameworks such as PyTorch, TensorFlow, ONNXRuntime, TensorRT, LLaMA.cpp, etc. • Proficiency in performance optimization for CPU, GPU, or AI hardware • Proficiency in kernel programming for accelerated hardware using programming models such as CUDA, OpenMP, OpenCL, Vulkan, and Metal • Experience with latency-sensitive optimizations and real-time inference • Understanding of resource constraints on mobile/edge hardware • Knowledge of model compression techniques (quantization, pruning, distillation, etc.)