AI Data Engineer
hoichoi is India's leading Bengali OTT platform, the digital arm of SVF, Eastern India's largest media & entertainment company. We stream 650+ Bengali movies, 185+ originals, and 3000+ hours of content to an audience of ~250 million Bengalis worldwide. We're part of the wider SVF ecosystem alongside Sooper (AI-led micro-drama OTT) and LoglineAI (AI-native creative studio), and we're on a mission to entertain everyone in their native language.
We are scaling our AI ambitions through a company-wide push to embed AI into how all we operate, anchored by a dedicated AI brain.
What you'll do :
hoichoi already runs a comprehensive data setup for consumer and content data. As AI at hoichoi expands, that same architecture needs to extend to cover it too — including turning unstructured content into knowledge that in-house AI agents and applications can actually use.
You'll own the day-to-day health of that data: schemas, availability, and fixing what breaks, working in close, daily contact with our engineering team and in direct support of the AI team.
- Process data requests from other teams — building views and data-set bundles for analytics use.
- Own day-to-day data hygiene: schema health, data availability, and fixing breakages as they occur.
- Work in close, daily contact with the engineering team; this is a high-involvement, high-contact role by design.
- Support the AI team with AI-ready data — including embeddings/vector pipelines, chunking, and handling unstructured content (video metadata, subtitles, synopses, thumbnails) — alongside your core data engineering work.
- Help extend hoichoi's existing data architecture to adopt new methodologies around transforming structured and unstructured data into knowledge consumed by in-house AI agents and applications as it grows.
- Build and maintain orchestration/transformation pipelines and streaming or event-ingestion pipelines (clickstream, CDC) that keep data flowing reliably.
- Own data governance and PII handling for consumer data, in line with DPDP Act compliance.
Must have qualifications :
- 4–8 years of experience in data engineering, with hands-on schema and pipeline work.
- Strong SQL fluency.
- Comfortable owning data hygiene end-to-end — not just building pipelines, but fixing what breaks in them.
- Experience with modern analytics/data technologies — ClickHouse, MongoDB, Postgres, or similar modern RDBMS/NoSQL databases.
- Experience with orchestration/transformation tooling (Airflow, Dagster, dbt, or equivalent) and streaming/event-ingestion pipelines (Kafka, CDC, clickstream).
- AI-first by default — you actively use AI coding tools (Cursor, Copilot, Claude Code, etc.) in your own work rather than defaulting to hand-coding everything, with at least one end-to-end example where AI generated significant code you debugged or overrode. This matters more to us than raw years of experience.
- Awareness of data governance, PII handling, and DPDP Act compliance — this role touches consumer data directly.
- Should stay on top of the latest Data & AI technology trends, and work closely with the AI and technology teams to bring relevant ones in.
Good to have :
- Experience in a media, streaming, or content platform environment.
- Exposure to preparing or structuring data specifically for AI/ML use cases.
- Excellent knowledge of modern data concepts — self-healing data pipelines, Knowledge Graphs, Ontologies.
- Data-quality/observability tooling experience (tests, monitoring, schema-drift alerting).
- Any relevant AI certifications (e.g., Anthropic Claude certifications) or data engineering/AI certifications from major hyperscalers.
Culture & How We Work
At hoichoi, the best idea wins, not the biggest title. We operate in small, cross-functional teams where content, product, growth, and tech sit close together, argue productively, and occasionally steal each other's snacks. It's a lot like a kitchen during a house party: slightly chaotic, everyone's doing three things at once, but somehow the food comes out great.
Ownership here isn't a buzzword; it's the default setting. If something's broken, you don't file a ticket and wait. You're the person who notices the printer is jammed and just fixes it instead of pretending you didn't see it. We ship fast, measure what works, and iterate. Think less 'perfect deck' and more 'let's see if real viewers actually care.'
When big moments arrive, we stretch. But we're equally serious about keeping the pace sustainable, because burnt-out teams build neither great products nor great stories. If that sounds like your kind of chaos, we should talk.
Recommended Jobs
Posted just now
Posted just now
Posted just now
Posted just now
Posted just now

