Website Humyn Labs
About Humyn Labs
Humyn Labs builds the intelligence layer for physical-world AI — systems that perceive, reason, and act in real environments. Our work sits at the intersection of egocentric video understanding, embodied AI, robotics perception, and voice-driven interaction. We move fast, obsess over data quality, and ship at scale.
Humyn Labs converts human action – across sound, sight, movement, and touch – into high-quality multi-modal data signals for physical AI. Operating across 20+ countries in India, southeast Asia, Latin America, and the Middle East: the real-world environments where physical AI deploys, not the labs where it is built.
Our data isn’t just collected; it’s evaluated, defended, and production-ready. Because before AI can be trusted, its training data must be.
Role Overview
We are looking for a Data Analyst to join Humyn Labs, turning complex, multi-source data into clear, decision-driving insight across speech AI benchmarking, dataset marketplace analytics, and platform reporting.
You will work closely with the Head of Data to produce benchmark analyses for IndicBench, build Superset dashboards, and support KAI (our NL2SQL tool) with robust SQL logic and schema documentation.
What You Will Work On
Speech AI Benchmarking
- Analyse ASR benchmark results across 16 languages and 14 providers for IndicBench, Humyn Labs’ speech recognition benchmarking initiative
- Interpret WER, CER, BERTScore, PIER, DER, and code-switching metrics
- Produce internal QA reports and publish-ready benchmark summaries that position Humyn Labs as a credible voice in speech AI evaluation
Dataset Marketplace & Platform Analytics
- Analyse usage, quality, and performance metrics across Humyn Labs’ dataset marketplace — tracking dataset uptake, buyer engagement, and quality signals across audio, image, and code modalities
- Support analysis tied to the Universal Data Schema (UDS) — helping monitor schema adoption, data consistency, and quality metrics across modalities
- Build reporting that helps the team understand where data quality issues originate and how they affect downstream marketplace credibility
KAI (NL2SQL) Support
- Validate NL2SQL outputs from Humyn Labs’ GPT-4-powered analytics tool against ground truth SQL
- Document schema, contribute FQA (Frequently Queried Analytics) examples, and maintain the Pinecone retrieval index
- Help make KAI a reliable self-serve analytics layer for non-technical stakeholders across the organization
Dashboards & Data Quality
- Build, maintain, and iterate on Superset dashboards for Humyn Labs leadership
- Resolve permission and connectivity issues; train team members on self-serve analytics
- Define and track data quality KPIs across ingestion pipelines; flag anomalies and escalate to engineering
Stakeholder & External Reporting
- Translate analytical findings into concise, credible outputs for founders, investors, and external partners
- Support LinkedIn research posts and benchmark publication materials that showcase Humyn Labs’ data and AI capabilities to the broader ecosystem
- Contribute directly to Humyn Labs’ credibility with dataset marketplace buyers through rigorous, well-documented analysis
You Must Have
- 2+ years in an analytical role with ownership of reporting and insight delivery
- Advanced SQL — window functions, CTEs, complex joins; comfortable writing production-quality queries against Athena or equivalent
- Proficiency with BI tooling — Superset, Metabase, Looker, or equivalent; able to build from scratch, not just edit
- Ability to work with Python for data wrangling (pandas, analysis notebooks) — you do not need to be an engineer, but you should be self-sufficient
- Exceptional written communication — you can write a benchmark report that is technically precise and externally publishable
- Comfort with ambiguity — Humyn Labs’ data landscape spans speech AI, dataset quality, and platform analytics; intellectual range matters
Strong Plus
- Experience with NLP evaluation metrics or ML model benchmarking
- Familiarity with dataset marketplaces, data licensing, or AI training data pipelines
- Prior experience presenting data findings to external stakeholders (investors, partners, or the public)
- Understanding of schema design principles (relevant to Humyn Labs’ Universal Data Schema work)
📌 How to Apply
👉 Apply now through DigitalSolutionTech.in to take the next step in your Data Analyst career and work with a mission-driven brand that blends science, technology, and empathy
To apply for this job email your details to hr@latestjobhiring.com
