AI ML Ops Engineer
About Tarento
Tarento is a fast-growing technology consulting company headquartered in Stockholm, with a strong presence in India and clients across the globe. We specialize in digital transformation, product engineering, and enterprise solutions, working across diverse industries including retail, manufacturing, and healthcare. Our teams combine Nordic values with Indian expertise to deliver innovative, scalable, and high-impact solutions.
We're proud to be recognized as a Great Place to Work, a testament to our inclusive culture, strong leadership, and commitment to employee well-being and growth. At Tarento, you’ll be part of a collaborative environment where ideas are valued, learning is continuous, and careers are built on passion and purpose.
Designation: Senior Software Engineer – Analytics
Location: Indore
Educational Qualifications: B.E/B.Tech
Exp: 3-5 Years
Mode of work: Hybrid
Role Overview
We are seeking an AI/ML Ops Engineer to own the deployment, monitoring, and scaling of ML models and pipelines in production. You will build the infrastructure and automation that lets models move reliably from training to production and stay healthy once there.
Key Responsibilities
Tarento is a fast-growing technology consulting company headquartered in Stockholm, with a strong presence in India and clients across the globe. We specialize in digital transformation, product engineering, and enterprise solutions, working across diverse industries including retail, manufacturing, and healthcare. Our teams combine Nordic values with Indian expertise to deliver innovative, scalable, and high-impact solutions.
We're proud to be recognized as a Great Place to Work, a testament to our inclusive culture, strong leadership, and commitment to employee well-being and growth. At Tarento, you’ll be part of a collaborative environment where ideas are valued, learning is continuous, and careers are built on passion and purpose.
Designation: Senior Software Engineer – Analytics
Location: Indore
Educational Qualifications: B.E/B.Tech
Exp: 3-5 Years
Mode of work: Hybrid
Role Overview
We are seeking an AI/ML Ops Engineer to own the deployment, monitoring, and scaling of ML models and pipelines in production. You will build the infrastructure and automation that lets models move reliably from training to production and stay healthy once there.
Key Responsibilities
- CI/CD for ML: Build and maintain automated pipelines for model training, testing, and deployment.
- Infrastructure: Containerize and orchestrate model-serving infrastructure (Docker, Kubernetes) at scale.
- Monitoring: Set up monitoring for model performance, data/concept drift, latency, and system health.
- Versioning & Reproducibility: Manage model/experiment versioning and reproducibility (MLflow, DVC, or similar).
- Cloud & GPU Management: Manage scalable, cost-efficient GPU/cloud infrastructure for training and inference workloads.
- Strong hands-on experience with Docker and Kubernetes
- CI/CD tooling (GitHub Actions, Jenkins, GitLab CI, or similar)
- Experience with model-serving frameworks (Triton Inference Server, TorchServe, or similar)
- Cloud platform experience (AWS/Azure/GCP), especially GPU infrastructure
- Python for automation and tooling
- Experience with experiment tracking and model registries (MLflow, DVC, Weights & Biases)
- Monitoring/observability tooling (Prometheus, Grafana, or similar)
- Production experience with Triton Inference Server: model repositories, ensembles, dynamic batching, instance groups
- GPU operations on Kubernetes: NVIDIA GPU Operator, MIG/time-slicing, node pools, driver/CUDA version management
- Hands-on with TensorRT / ONNX Runtime conversion and performance profiling (Nsight, perf_analyzer) gRPC and streaming service patterns; load testing tools (Locust, k6)
- Strong Linux, networking and debugging fundamentals for bare-metal environments
- Docker, Kubernetes, CI/CD, Python, cloud infrastructure (AWS/Azure/GCP)
- Experience deploying LLM/NLP inference services specifically (batching, quantization, low-latency serving)
- Familiarity with vLLM, Ollama, or similar for local LLM hosting
- Infrastructure-as-code experience (Terraform, Helm)
- Experience with NVIDIA NVCF, NeMo or NIM deployments
- Serving TTS/ASR models where time-to-first-byte matters
- Log/trace stacks (Loki, OpenTelemetry, ELK) and on-call tooling
Recommended Jobs
Digital Marketing Specialist
Posted just now
RETAIL TRAINEE PHARMACIST - Renowned Pharmacy, Nellore
Posted just now
Ab Initio Software Engineer
Posted just now
Intensivist
Posted just now
Content Developer
Posted just now

