Flam is building the next generation of interactive media through its content format. We are an AI-native technology company transforming how brands and consumers interact through immersive, interactive content. Our technology enables rich, app-less experiences that can be launched instantly on smartphones, creating a fundamentally different way for brands to engage consumers. We are backed by leading investors and already work with some of the world's largest brands. We are now building Flicks, our interactive media format for the US market.
About the role :
Every model we ship at Flam is bounded by the data that trained it, and every eval number we report is only as trustworthy as the eval set behind it. This role owns both. You will build and run the data layer underneath our LLM Falcon, our TTS system Finesse). That means text, audio, and video, cleaned, deduplicated,filtered,labelled, synthetically generated where real data doesn't exist, and versioned so that six months from now we can say exactly what went into a checkpoint. This is an engineering role. You will write Python every day. It is not an annotation or labelling-management position, though you will design annotation guidelines and quality-check what comes back.
What you'll do
Build ingestion and cleaning pipelines at scale — deduplication (exact, near-dup, semantic), quality filtering with classifier-based scoring, PII stripping,language identification, format normalization.
Generate synthetic data where real data is scarce or expensive: LLM generated instruction and preference data, code-mixed Indic text, rendered scenes for image models, augmented audio for speech models.
Curate multilingual and code-mixed Indic corpora. A large part of our differentiation is Indic-language quality, and a large part of that comes down
to data hygiene most pipelines get wrong.
Build and maintain evaluation sets. Design them so they measure what we think they measure, keep them uncontaminated, and version them properly
Handle multimodal data - audio segmentation and transcript alignment for TTS/ASR, face and video preprocessing for avatar training, image-caption pair curation.
Own data provenance and versioning. Which files, which filters, which version, which run. This should be answerable in one command, not one
afternoon. Write annotation guidelines and audit annotation quality when human labelling is in the loop.
What we're looking for:
Strong Python. You are comfortable processing datasets far larger than memory, and you know when to reach for a database instead of a script.
Practical data engineering fundamentals — streaming, chunking, parallelism, checkpointing long jobs, handling malformed input without losing the run.
Familiarity with the modern data formats and tooling of ML Parquet,WebDataset, HuggingFace Datasets, object storage S3/GCS/R2.
Genuine care about data quality. You should find it uncomfortable when a dataset has duplicates in it.
Enough ML understanding to know how a data decision propagates into model behaviour — why dedup matters, why eval contamination invalidates a benchmark, what a bad filter does to a distribution's tails.
Strongly preferred:
Experience curating training data for LLMs, speech models, or image/video models specifically.
Working knowledge of embeddings and vector search for semantic deduplication and retrieval FAISS, Milvus, or similar).
Native or near-native fluency in one or more Indian languages, with the ability to judge quality in it — this is a real asset here, not a checkbox.
Audio or video processing experience (ffmpeg, torchaudio, forced alignment).
Synthetic data generation using LLMs, or 3D rendering pipelines Blender, Unreal) for visual data.
Not required:
A degree in ML.
Prior model training experience.
Experience with our exact toolchain.
Why this role matters :
At most companies, data work is what gets handed to whoever is available. Here it's a named role with a named owner, because our model quality is directly downstream of it. You'll work alongside the engineers training the models, and you'll see your filtering decisions show up in benchmark numbers within the same quarter. If you want to move into model training over time, this is a strong path into it and we'll support that.