Writing.io Jobs

Find the best remote jobs. Answer a few questions and we'll deploy a powerful assistant to help you search, create alerts, and more.

1 What roles are you open to?

2 Experience level

3 Work style

Did you know? If memory is enabled, Writing.io can remember your job search preferences and help you to improve your resume, craft customized outreach and more.

Trainer ML Annotation QA Engineer at Gather AI

Analyzes and ensures quality of annotated training data for computer vision models, making judgment calls on annotation accuracy and identifying data drift or model regressions.

Mid Posted 2 days ago RemoteFirstJobs Product
What this role involves

Job Title: ML Annotation QA Engineer

About Us

Are you ready to build the future of supply chain? At Gather AI, we’re not just creating software, we’re pioneering a new era of warehouse intelligence. We’ve developed a groundbreaking, vision-powered platform that uses autonomous drones and existing equipment to capture real-time data, completely digitizing workflows that have historically been manual and error-prone. This means facilities operate smarter, safer, and more efficiently, ultimately redefining “on-time, in full” delivery.

If you’re looking for an opportunity to contribute to truly transformative technology and make a significant impact in a vital industry, Gather AI is the place for you. We’re leading the charge in the rapidly evolving robotics industry, and we invite you to join us in reshaping the global supply chain, one intelligent warehouse at a time.

About the Team

Our engineering organization spans autonomy, computer vision and machine learning, embedded and hardware systems, full-stack, and cloud, all working in parallel across multiple active product lines tied to live customer deployments. It’s a technically deep, fast-moving team where individual contributors carry real accountability and the work shows up directly in customer operations. Ground truth quality sits at the centre of that — the annotated data this role owns is what our models are trained and measured against.

About the Role

We are looking for an ML Annotation QA Engineer to own the quality of annotated data across our computer vision and machine learning programs. This role is responsible for the judgment-heavy analysis that cannot be reliably outsourced, for the decision rules behind it, and for turning annotation output into an ongoing read on how our systems are actually performing in the field.

We work with an external annotation partner at production volume, and that continues. What we need in house is someone who can analyse the annotated data, make and defend the calls the vendor cannot make consistently, and build the aggregate view that shows which facilities and equipment are degrading and why. Annotation drift, a model regression, a tool bug, and genuine field degradation all look similar in a chart and require completely different responses — telling them apart is the core of the job.

You will work closely with Machine Learning Engineers, QA, and Engineering, and the first assignment is our warehouse forklift vision program, where barcode readability and localisation analysis are the immediate need. From there the remit grows with us: drone imagery annotation today, and new task types as customer-driven capabilities come online. Success in this role requires a combination of analytical rigor, sound judgment under ambiguity, and clear written communication.

What You’ll Do

  • Own the judgment-heavy quality analysis on annotated data that cannot be reliably outsourced — working the daily review queue and producing verdicts and root cause in house.
  • Own, version, and refine the verdict taxonomy, decision rules, and quality guidelines for the categories you cover.
  • Build and maintain performance trackers over annotated data — error rates by facility, site, equipment, and data format over time, against an agreed baseline.
  • Detect anomalies against that baseline and flag them the day they appear rather than weeks later.
  • Run root cause analysis on flagged anomalies, distinguishing annotation error from model or system error from genuine degradation in the field.
  • Report findings to engineering and ML with reproducible evidence and stated confidence, fast enough that the issue is still observable.
  • Identify systematic failure patterns rather than one-off misses, and maintain a documented pattern library others can use.
  • Query and analyse annotation data directly with Python and SQL to test hypotheses, without waiting on extracts from anyone.
  • Feed annotation-quality findings back as concrete SOP and instruction changes when the root cause is labeling rather than system behaviour.
  • Specify annotation tool improvements — what the tool should surface so analysis stops requiring manual work — and validate the fixes.
  • Stand up quality analysis and reporting for new annotation programs as they come online.
  • Track work across Jira and contribute to pre-release validation for the behaviours you cover.

First 90 Days

Within your first three months, you will be expected to:

Own the Quality Analysis

  • Take over barcode and location root cause analysis from the annotation vendor.
  • Move from supervised review to owning the daily queue at the agreed review threshold, inside the expected time budget.
  • Become the primary owner of verdicts and root cause for the categories you cover.

Become Fluent in the Annotation Pipeline

Develop working expertise in the data and the domain behind it:

  • What is captured, at what grain, and where the known quality limits are
  • Racks, locations, levels, bins and bin types; LPN, SKU, AWB and other code formats; exceptions and exception types; OCR versus barcode and the failure modes of each

Build the Performance Tracker

  • Stand up a facility performance tracker with an agreed baseline, thresholds, and reporting cadence.
  • Establish anomaly detection against it, and get the team to the point of trusting and using it.
  • Investigate flagged anomalies independently, with a root cause hypothesis and stated confidence.

Own Documentation

Take ownership of:

  • Barcode and location verdict taxonomy and decision rules
  • The pattern library of known failure modes

Responsibilities include:

  • Version control
  • Closing documentation gaps
  • Adding edge-case guidance
  • Feeding changes back to the annotation vendor’s SOPs where the root cause is labeling

Required Technical Skills

  • BS in Computer Science/Engineering, Electrical Engineering, or equivalent experience
  • Experience working with annotated datasets for CV/ML, including assessing label quality
  • Strong understanding of statistics, and able to work with data
  • Root cause analysis — generating competing hypotheses and naming the evidence that separates them
  • Familiarity with Python and SQL, or equivalent, for querying and analysing data independently
  • Writing quality guidelines, decision rules, and labeling taxonomies
  • Excellent documentation, communication, and collaboration skills

Nice-to-Have Skills

  • 2+ years in quality assurance, data quality, or product support
  • 1+ years experience with enterprise-grade ticketing systems (e.g. Jira)
  • Computer vision annotation experience as a reviewer or auditor — video event labeling, bounding boxes, polygon segmentation, counting, or classification
  • Annotation quality methodology — gold-set validation, inter-annotator agreement, sampling design
  • Strong spatial and geometric reasoning — relevant wherever labels describe position in physical space
  • Experience in warehouse automation, robotics, or computer vision applications
  • Dashboarding or BI tooling for recurring reports
  • Small team experience

Qualifications

Required

  • BS in Computer Science/Engineering, Electrical Engineering, or equivalent experience
  • 2–5 years of experience in ML QA, annotation quality, data quality, or analytics at an AI/ML company
  • Hands-on experience with annotated ML datasets, including assessing label quality
  • Demonstrated experience finding, diagnosing, and reporting data anomalies to a technical audience
  • Able to get to a defensible answer from unfamiliar data without someone preparing it first
  • Understanding of data privacy and confidentiality requirements when working with customer operational data

Preferred

  • Experience in warehouse automation, robotics, or computer vision applications
  • Experience with annotation platforms and quality tooling
  • Experience authoring annotation guidelines or standard operating procedures

Required Technologies

  • Python
  • SQL, or an equivalent query language
  • Jira

Success Traits

We’re looking for someone who is:

  • Analytical, able to tell a real pattern from noise and to say honestly when the data will not settle a question
  • Detail-oriented — a verdict or a tracker that is subtly wrong is worse than none at all
  • Comfortable with ambiguity, and willing to present competing hypotheses rather than forcing a single answer
  • Persistent in chasing a cause across sites, programs, and builds
  • Resourceful and comfortable building processes and reports where they don’t yet exist
  • A strong written communicator — the output of this role is reports other people act on
  • Collaborative, respectful, and an excellent cross-functional partner

What Success Looks Like

By the end of your first six months, you will have:

  • Established yourself as the primary owner of annotation quality analysis and root cause
  • A performance tracker the team relies on, running on a regular cadence
  • Closed the reporting cycle to same day, so issues are raised while still observable
  • Standardized the verdict taxonomy and decision rules for the categories you cover
  • Built a documented pattern library of known failure modes that others can use
  • Improved the annotation tool in at least one way that measurably reduces manual effort
  • Given ML and engineering a clearer picture of what error patterns mean for model accuracy and field performance
  • Created a repeatable method for standing up annotation quality on the next program
Read the full description
Trainer Alpheratz Project - French (France) Translation Quality Reviewer

Reviews and rates French customer service content translations to ensure quality and accuracy for AI training datasets.

Mid Posted 8 days ago Himalayas
What this role involves
We are looking for experienced language professionals to support an ongoing project focused on reviewing and refining French (France) customer service content.
Read the full description
Trainer Software Engineer - AI Agent Evaluation (Remote) at Mindrift

Create and evaluate coding tasks for AI agents by building realistic developer environments, designing challenges, and writing tests to assess model performance.

Mid Remote Posted 10 days ago RemoteFirstJobs Product
What this role involves

Please submit your CV in English and indicate your level of English proficiency.

Mindrift connects specialists with project-based AI opportunities for leading tech companies, focused on testing, evaluating, and improving AI systems. Participation is project-based, not permanent employment.

What this opportunity involves We’re building a dataset to evaluate AI coding agents - how well a model handles real-world developer tasks.

You’ll create challenging tasks and evaluation criteria within realistic simulated environments:

  • Build realistic developer environments - a virtual company with codebase, infrastructure, and context (tickets, docs, conversations) that forms a believable development history
  • Design tasks from intermediate states of these environments - craft the prompt, define what “solved” means, and ensure the task is solvable by an AI agent
  • Write tests that verify agent solutions - accept all valid approaches and reject incorrect ones, neither too strict nor too lenient
  • Iterate on tasks and tests based on QA feedback - review agent solutions, analyze failures, and refine until the evaluation is fair and robust

What this is NOT

  • Not data labeling
  • Not prompt engineering
  • Not writing code from scratch - the agent writes most of the code; you guide and evaluate

What we look for

  • 5+ years in software development
  • Core stack: Python (FastAPI), JavaScript/TypeScript (React), Docker, Postgres, Kafka, Redis
  • Experience writing tests (functional, integration)
  • English proficiency - B2+

Why this is hard

Frontier models are already good at coding. Creating a task that genuinely challenges the best models is non-trivial. You need to deeply understand where models fail and what scenarios reveal the difference between a good and a bad solution. Tasks have many valid solutions - writing tests that accept all correct solutions and reject incorrect ones is harder than it sounds.

How it works

Apply → Pass qualification(s) → Join a project → Complete tasks → Get paid

Compensation

Up to $50/hr equivalent, depending on level and pace. Tasks are estimated at ~20 hours each; you set your own schedule.

Read the full description
Trainer Content moderator - EdTech at Securly

Review AI-flagged student online activity for safety risks, conduct risk assessments, and alert school personnel to potential harm situations in real-time.

Mid Remote Posted 10 days ago RemoteFirstJobs Product
What this role involves

Role Overview & Mission: Securly On-Call

As a member of the Securly On-Call team (formerly Securly 24), you are the critical human intelligence that powers our safety mission. Reporting to the Director of Student Safety, you play a vital role in reviewing online activity recognized and “flagged” by Securly’s AI as potentially harmful to students. You will work in tandem with our award-winning technology to perform thorough risk assessments, distinguishing between curiosity and crisis, and executing critical communication protocols to alert school safety teams. This is a high-impact, remote-first role requiring a mission-driven mindset to support student well-being in real-time.

  • Location: 100% fully Remote (candidates must live in US or UK based)
  • Shift Hours: Monday - Friday 7-3 Central US Time (some flexibility possible)

Impact of the Securly On-Call Team:

  • Life-Saving Response: Our analysts notify school personnel of extreme risk situations—including potential self-harm, suicide, or violence—within 5 min or less.
  • Beyond the AI: You will conduct deep-dive analyses of student activity history to provide schools with the critical context needed for effective, life-saving interventions.
  • 24⁄7 Protection: Safety never sleeps. Our team provides a watchful eye whenever schools need us most, whether during the school day, after hours, or 24/7/365.
  • See Your Impact: Watch this Playlist of Analyst Testimonials to hear directly from our team about the profound impact this role has on student lives.

About Securly

Securly is the market leader in AI-powered student wellness and safety solutions, protecting over 20M students across 20,000+ schools worldwide. Our technology has analyzed over 10B activities to help keep students safe, secure, and ready to learn. We operate with the scale of a tech giant and the agility of a startup, combining innovation, data, and compassion to solve real problems.

Our Impact & Recognition

  • 2,000+ student lives saved.
  • Named to the 2026 GSV 150, honoring the world’s top 150 companies driving growth and innovation in digital learning.
  • Repeatedly recognized as a Top Place to Work and for EdTech Product of the Year.
  • Named to TIME’s World’s Top EdTech Companies of 2026

Our Culture of Engagement

At Securly, culture is our foundation. We are a fully remote, people-first organization built on trust and accountability. Our engagement data far exceeds global benchmarks:

  • 82% employee engagement (vs. 73% global).
  • 94% of employees are proud to work for Securly (vs. 81% global).
  • 91% rate manager effectiveness as high (vs. 77% global).

Key Expectations in First 12 Months

  1. High-Impact AI Support: Support Securly’s student safety AI suite by accurately reviewing and categorizing potentially harmful online activity.
  2. Safety Protocol Mastery: Execute complex communication processes with the Student Safety Group to proactively prevent harm and promote student well-being.
  3. Queue & Metric Excellence: Maintain consistent ownership of the online alert queue, meeting rigorous goals for response time, documentation accuracy, and resolution quality.

Key Development Milestones

  • 30 Days: Complete comprehensive training on Securly’s AI safety suite, emergency communication protocols, and internal support software.
  • 60 Days: Independently manage the live alert queue while maintaining high accuracy in identifying harmful vs. non-harmful content.
  • 90 Days: Master all emergency escalation workflows; independently manage communications with school districts via email and phone.
  • 6 Months: Identify recurring student behavior patterns or emerging digital trends and contribute improvements to internal safety documentation.
  • 12 Months: Serve as a trusted subject matter expert; consistently meet or exceed expectations and provide insights to evolve safety technology.

Key Responsibilities

  • Incident Analysis: Review potentially harmful online activity flagged by AI and decide on appropriate intervention levels.
  • Crisis Response & Communication: Manage emergency protocols and communicate with school districts via email or phone to prevent harm.
  • Operational Queue Management: Actively monitor the real-time online queue to ensure no critical incident is missed.
  • Data & Documentation: Maintain meticulous records in support software, detailing investigative steps and outcomes for all incidents.

Qualifications & Skills

  • Applying Domain Expertise: Leverage a background in education, psychology, or law enforcement to accurately distinguish between typical student behavior and high-stakes crisis indicators.
  • Executing Under Pressure: Draw upon past crisis intervention or counseling experience to execute immediate, high-pressure communication protocols during extreme-risk incidents.
  • Analytical Endurance: Maintain high-level concentration to manage a continuous real-time queue, ensuring all critical alerts are documented and escalated within the required 5 min window.
  • Translating Culture: Regularly interpret social media trends and pop culture slang to provide schools with the critical context needed for effective safety interventions.
  • Bilingual Intervention: Utilize bilingual proficiency to conduct real-time safety

Why Join Securly

  • Direct Mission Impact: Join the GSV 150–recognized team that is the direct “human intelligence” behind student life-saving interventions.
  • Industry-Leading Technology: Work with the most innovative, AI-powered safety suite trusted by 20,000+ districts worldwide.
  • Remote-First, People-First: Thrive in a trust-based culture with high engagement and 91% manager effectiveness ratings.
  • Professional Mastery: Build deep expertise in digital student safety and crisis management within a supportive, mission-driven team.

Wellness & Benefits Overview

  • Work-Life Balance: Remote-first work culture with Unlimited Vacation (Flex Time), 13 paid holidays, and a full week of paid holiday break around Christmas and New Year.
  • Parental & Family Support: 12 weeks of fully paid parental leave for birth, adoption, or fostering.
  • Health & Well-being: Company-sponsored medical, dental, and vision benefits, comprehensive EAP, and unlimited free access to well-being and mental health resources.
  • Financial Security: 401(k) plan with employer match and tax-advantaged spending accounts (HSA/FSA).
  • Professional Growth: $1,000 annual stipend for professional development to support continuous learning.
  • Perks: Exclusive access to the PerkSpend discount platform for savings on travel, electronics, gym memberships, and more. Enjoy a culture recognized as a Top Place to Work for 3+ years running.

Equal Opportunity Employer

Securly is committed to building a diverse and inclusive workplace. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, disability, or any other legally protected characteristic. Accommodations are available throughout the hiring process. Please contact recruitment.us@securly.com.

#LI-REMOTE #LI-DO1

Read the full description
Trainer Avid Media Composer Editor

Uses Avid Media Composer to annotate and label video data for AI model training and development.

Mid Remote Posted 18 days ago Himalayas
What this role involves
Role Title: Avid Media Composer EditorRole Type: ContractorLocation: Remotemicro1 is engaging Avid Media Composer Editors to support a customer's project in high-impact video data annotation and AI model training.
Read the full description
Trainer TrainBrainSpace: Expert AI Trainers in Coding, Finance & Accounting, and Medicine/Nursing

Expert practitioner writes, solves, and grades training tasks for AI models in your field of expertise, providing structured feedback on model outputs.

Mid Remote Posted 19 days ago We Work Remotely — Programming
What this role involves

Headquarters: Atlanta, AZ.
URL: https://trainbrain.space/

What if the work you already do every day — shipping code, closing books, reading charts — became the training signal that teaches frontier AI how to think in your field?

TrainBrain builds expert-authored training data and evaluations for AI labs. We're hiring practising professionals to write, solve, and grade the tasks advanced models are measured against. When a model returns a broken function, a wrong reconciliation, or an unsafe clinical suggestion, it's because no expert caught it during training. You'd be that expert.

This isn't your day job in a new wrapper. There are no clients, no month-end close, no tickets, no on-call. You work in writing, on your own schedule, on problems chosen to sit right at the edge of what current models can handle.

No AI background required — we train you on the tooling.

Track 1 — Software Engineering

Write, debug, and optimise code across languages and problem domains. Review AI-generated solutions for correctness, security, and performance, and design edge cases that expose model weaknesses. Provide structured feedback explaining why code works — or doesn't — and compare competing solutions on engineering quality.

You'll need:

  • Advanced proficiency in Python, JavaScript/TypeScript, Java, C/C++, Go, or Rust
  • A solid grasp of data structures, algorithms, and software design principles
  • A CS degree or equivalent professional experience

Nice to have:

  • Open-source contributions or competitive programming
  • Distributed systems, cloud, or DevOps background
  • Code review, technical writing, or mentoring experience

Track 2 — Finance & Accounting

Author reconciliation scenarios: payment-to-invoice matching, billing versus recognised revenue, margin and variance analysis. Build the underlying data and source documents so each scenario is realistic, write rubrics specifying exact figures and required reasoning steps, then solve every task yourself to confirm the numbers hold.

You'll need:

  • Hands-on experience as an accountant, financial analyst, controller, or auditor
  • Working knowledge of revenue recognition, matching principles, and standard close workflows
  • Comfort with financial documents, spreadsheets, and accounting system outputs

Nice to have:

  • CPA, CA, ACCA, CIMA, or CFA
  • IFRS or US GAAP depth
  • ERP familiarity (SAP, Oracle, NetSuite, QuickBooks)
  • Audit, financial controls, or revenue assurance background

Track 3 — Medical & Clinical

Challenge models on differential diagnosis, drug interactions, treatment protocols, pathophysiology, and evidence appraisal. Write clinical vignettes with defensible, sourced answers, verify model outputs against current evidence, and document precisely where reasoning breaks down.

You'll need:

  • An MD, DO, MBBS, PharmD, RN/BSN, or advanced health sciences degree
  • Real clinical or research experience
  • Command of medical terminology and clinical reasoning

Nice to have:

  • Board certification or active licensure
  • A specialty focus (internal medicine, oncology, psychiatry, emergency medicine, radiology, nursing practice)
  • Peer-reviewed publications
  • Clinical trial, epidemiology, or medical education experience

What the work demands

Whatever your track, four things matter more than anything else.

Show your work. Making your reasoning explicit matters as much as the answer itself.

Be verifiable. Every task you write needs a correct answer someone else can confirm.

Iterate. You'll refine tasks until difficulty and gradability are both right.

Flag uncertainty. Saying "I'm not sure, and here's why" is a feature, not a failure.

You'll also need excellent written English, sharp attention to detail, and a secure computer with reliable internet.

Terms

This is a 1099 independent contractor position, not W-2 employment. You're responsible for your own taxes, and company-sponsored benefits such as health insurance, PTO, and retirement contributions don't apply. Contractors outside the US are engaged under the equivalent local arrangement.

Pay: $45–$100/hr, set by domain, seniority, and assessment performance. Specialised and hard-to-source expertise sits at the top of the band, and you're paid on a regular cadence.

Schedule: Choose your own projects, hours, and volume. Scale up in quiet weeks, scale down when your day job gets busy.

Growth: Strong contributors are invited into higher-rate specialist projects and review roles, with ongoing work as new engagements launch.

To apply: https://weworkremotely.com/remote-jobs/trainbrainspace-expert-ai-trainers-in-coding-finance-accounting-and-medicine-nursing

Read the full description
Trainer Ex-MBB Strategy Consultant - AI Training (Remote) at Mindrift

Ex-MBB consultants create structured learning environments and tasks to train AI models on real-world consulting problem-solving and business reasoning.

Mid Remote Posted 21 days ago RemoteFirstJobs Product
What this role involves

Toloka AI supports frontier model post-training by building domain-specific reinforcement learning environments, tasks, and evaluation frameworks designed by real practitioners.

Mindrift, powered by Toloka — a leading enterprise AI and machine learning data partner since 2014 — connects top domain experts with cutting-edge AI initiatives. Backed by Toloka’s deep expertise in scalable data generation, crowd technology, and applied ML systems, Mindrift enables experts to shape how next-generation generative models learn, reason, and perform.

We are launching a Management Consulting domain focused on translating real-world consulting engagements into structured learning environments for advanced AI systems. To do this credibly, we are assembling a team of strategy consultants from top-tier firms who can convert authentic project experience into end-to-end examples — from problem structuring and work planning to analysis, synthesis, and client-ready recommendations.

You will join a growing team of consultants from leading strategy firms shaping how AI learns high-level business reasoning.

Important: This role is exclusively for consultants with direct experience at a top-tier strategy consulting firm. If you do not have hands-on project experience at one of the firms listed below, please do not apply. This requirement ensures the domain is built by practitioners trained to the highest standards of structured problem-solving and client delivery.

Eligible firms: McKinsey & Company, Boston Consulting Group (BCG), Bain & Company, Oliver Wyman, Roland Berger, Monitor Deloitte (Deloitte S&C), EY-Parthenon, Kearney, and Strategy& (PwC).

Who We’re Looking For

Consultants with 3+ years of experience at one of the firms listed above, with hands-on project experience in:

  • Structuring ambiguous client problems into workable analytical plans
  • Building financial models, market analyses, or synthesized findings from messy inputs
  • Producing client-ready deliverables under time pressure
  • Forming and defending recommendations under uncertainty

No deep technical background is required — we will onboard you on the lightweight tools involved.

What You’ll Do

  • Build realistic consulting project environments — create detailed project scenarios grounded in real engagement dynamics: industry context, financials, constraints, conflicting inputs, and incomplete information.
  • Design structured consulting tasks for AI agents— break projects into discrete tasks that mirror real consulting work: market sizing, commercial due diligence, cost optimization, growth strategy, operational diagnosis, benchmarking, and more.
  • Define evaluation criteria and quality standards — develop grading frameworks, evaluation rubrics, and golden-answer solutions for each task, used to train and calibrate an LLM-based grading system that evaluates AI outputs at scale.

This is a remote, project-based, individual-contributor role focused on analytical design and evaluation.

Skills & Requirements

  • 3+ years at McKinsey, BCG, Bain, Oliver Wyman, Roland Berger, Monitor Deloitte, EY-Parthenon, Kearney, or Strategy&
  • Strong structured problem-solving and hypothesis-driven thinking
  • Ability to translate vague problems into clear analytical steps and deliverables
  • High attention to logical consistency and output quality
  • Independent, self-directed working style
  • Clear written English (B2+)

Compensation

On this project, contributors can earn up to $60 per hour equivalent, depending on their level and pace of contribution.

Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.

For this project, tasks are estimated to require around 25-30 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Read the full description
Trainer Ex-MBB Strategy Consultant - AI Training (Remote) at Mindrift

Ex-MBB consultant creates realistic consulting project scenarios and structured learning tasks to train AI systems on business reasoning and problem-solving.

Mid Remote Posted 21 days ago RemoteFirstJobs Product
What this role involves

Toloka AI supports frontier model post-training by building domain-specific reinforcement learning environments, tasks, and evaluation frameworks designed by real practitioners.

Mindrift, powered by Toloka — a leading enterprise AI and machine learning data partner since 2014 — connects top domain experts with cutting-edge AI initiatives. Backed by Toloka’s deep expertise in scalable data generation, crowd technology, and applied ML systems, Mindrift enables experts to shape how next-generation generative models learn, reason, and perform.

We are launching a Management Consulting domain focused on translating real-world consulting engagements into structured learning environments for advanced AI systems. To do this credibly, we are assembling a team of strategy consultants from top-tier firms who can convert authentic project experience into end-to-end examples — from problem structuring and work planning to analysis, synthesis, and client-ready recommendations.

You will join a growing team of consultants from leading strategy firms shaping how AI learns high-level business reasoning.

Important: This role is exclusively for consultants with direct experience at a top-tier strategy consulting firm. If you do not have hands-on project experience at one of the firms listed below, please do not apply. This requirement ensures the domain is built by practitioners trained to the highest standards of structured problem-solving and client delivery.

Eligible firms: McKinsey & Company, Boston Consulting Group (BCG), Bain & Company, Oliver Wyman, Roland Berger, Monitor Deloitte (Deloitte S&C), EY-Parthenon, Kearney, and Strategy& (PwC).

Who We’re Looking For

Consultants with 3+ years of experience at one of the firms listed above, with hands-on project experience in:

  • Structuring ambiguous client problems into workable analytical plans
  • Building financial models, market analyses, or synthesized findings from messy inputs
  • Producing client-ready deliverables under time pressure
  • Forming and defending recommendations under uncertainty

No deep technical background is required — we will onboard you on the lightweight tools involved.

What You’ll Do

  • Build realistic consulting project environments — create detailed project scenarios grounded in real engagement dynamics: industry context, financials, constraints, conflicting inputs, and incomplete information.
  • Design structured consulting tasks for AI agents— break projects into discrete tasks that mirror real consulting work: market sizing, commercial due diligence, cost optimization, growth strategy, operational diagnosis, benchmarking, and more.
  • Define evaluation criteria and quality standards — develop grading frameworks, evaluation rubrics, and golden-answer solutions for each task, used to train and calibrate an LLM-based grading system that evaluates AI outputs at scale.

This is a remote, project-based, individual-contributor role focused on analytical design and evaluation.

Skills & Requirements

  • 3+ years at McKinsey, BCG, Bain, Oliver Wyman, Roland Berger, Monitor Deloitte, EY-Parthenon, Kearney, or Strategy&
  • Strong structured problem-solving and hypothesis-driven thinking
  • Ability to translate vague problems into clear analytical steps and deliverables
  • High attention to logical consistency and output quality
  • Independent, self-directed working style
  • Clear written English (B2+)

Compensation

On this project, contributors can earn up to $60 per hour equivalent, depending on their level and pace of contribution.

Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.

For this project, tasks are estimated to require around 25-30 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Read the full description
Trainer Ex-MBB Strategy Consultant - AI Training (Remote) at Mindrift

Ex-MBB consultant creates realistic consulting project scenarios and structured tasks to train AI models on business reasoning and problem-solving.

Mid Remote Posted 21 days ago RemoteFirstJobs Product
What this role involves

Toloka AI supports frontier model post-training by building domain-specific reinforcement learning environments, tasks, and evaluation frameworks designed by real practitioners.

Mindrift, powered by Toloka — a leading enterprise AI and machine learning data partner since 2014 — connects top domain experts with cutting-edge AI initiatives. Backed by Toloka’s deep expertise in scalable data generation, crowd technology, and applied ML systems, Mindrift enables experts to shape how next-generation generative models learn, reason, and perform.

We are launching a Management Consulting domain focused on translating real-world consulting engagements into structured learning environments for advanced AI systems. To do this credibly, we are assembling a team of strategy consultants from top-tier firms who can convert authentic project experience into end-to-end examples — from problem structuring and work planning to analysis, synthesis, and client-ready recommendations.

You will join a growing team of consultants from leading strategy firms shaping how AI learns high-level business reasoning.

Important: This role is exclusively for consultants with direct experience at a top-tier strategy consulting firm. If you do not have hands-on project experience at one of the firms listed below, please do not apply. This requirement ensures the domain is built by practitioners trained to the highest standards of structured problem-solving and client delivery.

Eligible firms: McKinsey & Company, Boston Consulting Group (BCG), Bain & Company, Oliver Wyman, Roland Berger, Monitor Deloitte (Deloitte S&C), EY-Parthenon, Kearney, and Strategy& (PwC).

Who We’re Looking For

Consultants with 3+ years of experience at one of the firms listed above, with hands-on project experience in:

  • Structuring ambiguous client problems into workable analytical plans
  • Building financial models, market analyses, or synthesized findings from messy inputs
  • Producing client-ready deliverables under time pressure
  • Forming and defending recommendations under uncertainty

No deep technical background is required — we will onboard you on the lightweight tools involved.

What You’ll Do

  • Build realistic consulting project environments — create detailed project scenarios grounded in real engagement dynamics: industry context, financials, constraints, conflicting inputs, and incomplete information.
  • Design structured consulting tasks for AI agents— break projects into discrete tasks that mirror real consulting work: market sizing, commercial due diligence, cost optimization, growth strategy, operational diagnosis, benchmarking, and more.
  • Define evaluation criteria and quality standards — develop grading frameworks, evaluation rubrics, and golden-answer solutions for each task, used to train and calibrate an LLM-based grading system that evaluates AI outputs at scale.

This is a remote, project-based, individual-contributor role focused on analytical design and evaluation.

Skills & Requirements

  • 3+ years at McKinsey, BCG, Bain, Oliver Wyman, Roland Berger, Monitor Deloitte, EY-Parthenon, Kearney, or Strategy&
  • Strong structured problem-solving and hypothesis-driven thinking
  • Ability to translate vague problems into clear analytical steps and deliverables
  • High attention to logical consistency and output quality
  • Independent, self-directed working style
  • Clear written English (B2+)

Compensation

On this project, contributors can earn up to $60 per hour equivalent, depending on their level and pace of contribution.

Compensation varies across projects depending on scope, complexity, and required expertise. Please note that other projects on the platform may offer different earning levels based on their requirements.

For this project, tasks are estimated to require around 25-30 hours per week during active phases, based on project requirements. This is an estimate, not a guaranteed workload, and applies only while the project is active. Tasks must be submitted by the deadline and meet the listed acceptance criteria to be accepted.

Read the full description
Trainer Expert Audio Transcriber, Egyptian Arabic

Transcribes audio in Egyptian Arabic to create high-quality training data for AI models.

Mid Remote Posted 21 days ago Himalayas
What this role involves
ABOUT PERLE Perle is an AI infrastructure company building expert-driven training data, evaluation systems, and applied AI products for the world's leading labs and enterprises.
Read the full description
Trainer Expert Audio Transcriber, Portuguese (Brazilian)

Transcribes audio in Brazilian Portuguese to create high-quality training data for AI models.

Mid Remote Posted 21 days ago Himalayas
What this role involves
ABOUT PERLE Perle is an AI infrastructure company building expert-driven training data, evaluation systems, and applied AI products for the world's leading labs and enterprises.
Read the full description
Trainer Garment Manufacturing QC Specialist – Freelance AI Trainer Project

Applies garment manufacturing QC expertise to train AI models by providing feedback and labeling data for quality control automation systems.

Mid Remote Posted 24 days ago Jobicy AI
What this role involves
Are you experienced in garment manufacturing quality control and interested in helping train the next generation of AI systems? As AI models are increasingly applied to global manufacturing, supply chain...
Read the full description
Trainer Toptal: [Expert Crowd] Personal Finance consultants for AI model

Personal finance expert evaluates and trains AI models by providing feedback on financial advice scenarios using real-world expertise.

Mid Remote Posted 27 days ago We Work Remotely — Programming
What this role involves

Headquarters:

We are looking for experienced personal finance professionals to support the training and evaluation of AI models. You should have hands-on experience advising individual clients or households on real-world financial decisions.

Areas of Expertise

  • We are seeking talent with experience in one or more of the following:
  • Personal Tax: Individual tax returns (1040), tax planning, capital gains, self-employment and quarterly taxes
  • Investments & Retirement: Portfolio management, asset allocation, 401(k), IRA/Roth, rebalancing and Social Security
  • Financial Planning: Budgeting, cash flow, debt, financial goals and comprehensive household planning
  • Estate Planning: Estate/gift tax, trusts, beneficiary planning and fiduciary accounting (non-attorney roles)
  • Debt & Life Events: Student loans, mortgages, credit cards, home affordability and 529/college planning
  • Equity Compensation & Self-Employed Finances: RSUs, stock options, ESPPs and freelancer/self-employed finances

Ideal Candidates

  • US-based professionals with direct experience advising individuals or households
  • 5+ years of relevant experience preferred
  • Credentials such as CFP, CPA, EA, PFS or CTFA are a plus
  • Comfortable explaining financial concepts and evaluating real-world personal finance scenarios
  • Experience working with everyday/middle-income households is especially valuable

To apply: https://weworkremotely.com/remote-jobs/toptal-expert-crowd-personal-finance-consultants-for-ai-model

Read the full description
Trainer AI Red Team Specialist - Remote | Upto $22/hr

Tests AI systems for vulnerabilities and biases by attempting to break or misuse models to improve their safety and robustness.

Mid Remote Posted 27 days ago Himalayas
What this role involves
About the jobMercor connects elite creative and technical talent with leading AI research labs.
Read the full description