Ten AI data services staffed from Latin America, each with defined deliverables, a quality metric written into the contract, and a bilingual team working your hours. Pick one service or run several in the same pod.
RLHF & Preference Optimization
Pairwise comparisons, rankings and critique loops that turn human judgement into reward-model training data for PPO and DPO pipelines.
Deliverables
- Pairwise preference datasets with rationale fields
- Multi-turn dialogue rankings
- Reward model training splits + held-out evals
QA: Inter-annotator agreement κ ≥ 0.80$26–36/hr
Hire rlhf & preference optimization specialists
Supervised Fine-Tuning (SFT) Data
Instruction-response pair creation, taxonomy tagging and rubric execution for domain-specific fine-tuning in English, Spanish and Portuguese.
Deliverables
- Instruction-response pairs with golden-set calibration
- Domain taxonomies and labeling guidelines
- Rubric scoring with dual review on high-stakes items
QA: Label accuracy ≥ 95% on golden sets$23–33/hr
Hire supervised fine-tuning (sft) data specialists
Data Collection & Creation
Net-new human-generated data: prompts, long-form writing, speech recordings, screen and device capture, and scenario scripting for edge cases your logs never contain.
Deliverables
- Prompt and response corpora written to spec
- Audio, image and video capture with consent records
- Scenario libraries for rare and adversarial cases
QA: Spec compliance ≥ 97% at acceptance sampling$18–28/hr
Hire data collection & creation specialists
Model Safety & Red Teaming
Adversarial testing of deployed and pre-release models: jailbreaks, prompt injection, harmful content probing and agentic misuse scenarios.
Deliverables
- Attack taxonomies with reproducible prompts
- Severity-scored findings and mitigation notes
- Regression suites for each patched failure mode
QA: Attack coverage across 12+ harm categories$28–40/hr
Hire model safety & red teaming specialists
Model Evaluation & Benchmarking
Human evals, LLM-as-judge calibration, golden datasets and regression suites so every prompt or model change is measured instead of guessed.
Deliverables
- Golden datasets with versioning
- Human eval rubrics and calibration sets
- Release-gate scorecards per model version
QA: Judge-human correlation tracked per release$25–35/hr
Hire model evaluation & benchmarking specialists
Multilingual & Localization Data
Native Spanish, Portuguese and English data creation, translation review, and locale-aware safety labeling for models serving the Americas.
Deliverables
- Native-quality ES/PT/EN datasets
- Locale-specific safety and toxicity labels
- Cultural adequacy review of model outputs
QA: Native reviewer sign-off on 100% of batches$20–30/hr
Hire multilingual & localization data specialists
Code & Agentic Traces
SWE-bench style task annotation, tool-calling traces, UI intents and agent trajectories labeled by engineers who read and run the code.
Deliverables
- Repository-level task annotations with test validation
- Tool-invocation and function-calling traces
- Agent trajectory scoring with failure taxonomies
QA: Functional correctness verified by test runs$30–45/hr
Hire code & agentic traces specialists
Multimodal Annotation
Vision, audio and video labeling: bounding boxes, segmentation, transcription, diarization and cross-modal alignment checks.
Deliverables
- Bounding boxes, polygons and segmentation masks
- Transcription and speaker diarization
- Cross-modal alignment verification
QA: Annotation precision audited on blind golden sets$18–30/hr
Hire multimodal annotation specialists
RL Environments
Task environments, simulators and verifiable reward functions for training and evaluating agents on real software workflows.
Deliverables
- Containerized task environments with deterministic setup
- Verifiable reward functions and scoring harnesses
- Difficulty-graded task suites with reference solutions
QA: Deterministic reproduction on 100% of tasks$32–48/hr
Hire rl environments specialists
Robotics & Teleoperation Data
Managed teleoperation sessions, manipulation trajectories and force/torque traces for training and evaluating real-world robot policies.
Deliverables
- Teleoperated demonstrations with observation-action pairs
- Episode success scoring and failure taxonomies
- Sensor-synced trajectories delivered as HDF5 or RLDS
QA: Deterministic replay checks on every episode$28–42/hr
Hire robotics & teleoperation data specialists
Egocentric Video & World Model Data
First-person capture with depth, segmentation, hand pose and action boundaries — the interaction video world models and VLA systems cannot scrape.
Deliverables
- Wearable-camera capture with consent records
- Depth, segmentation and hand-pose enrichment
- Action-boundary and affordance annotation
QA: Annotation agreement κ ≥ 0.80 on action boundaries$20–30/hr
Hire egocentric video & world model data specialists
Domain Expert Networks
Licensed and credentialed professionals — clinicians, lawyers, accountants, engineers — writing and reviewing data where a generalist annotator cannot.
Deliverables
- Expert-written Q&A and reasoning traces
- Expert adjudication of contested labels
- Domain rubric design and annotator training
QA: Credential verification on every expert$35–70/hr
Hire domain expert networks specialists