Draft · Inaugural season

Rules, Tracks & Scoring

How the Humanity Games are meant to work: who can enter, how challenge tracks are chosen, how teams combining Human ingenuity, AI intelligence, and Robotic precision are scored, what counts as a real-world outcome, and how the jury is selected.

Draft for discussion. This page is a draft for the inaugural season of Humanity Games. Everything here is subject to change and to legal review and does not yet form a binding agreement. No dates, fees, or prizes have been set.

Legend: a small amber code such as D1 marks an open decision — tap it to jump to the Open decisions list at the end of that section · Legal review marks text that needs counsel before publication.

Section 1

Official Rules (draft)

Plain-language rules for the inaugural season. Where something is not settled yet, we say so openly.

1. Purpose

Humanity Games is a global competition where teams combine Human ingenuity, AI intelligence, and Robotic precision (“H-AI-R” collaboration) to make real progress on global challenges — and share that journey openly so others can learn from it. The goal is useful, safe, demonstrable solutions, not hype.

2. Eligibility

  • Open to individuals from anywhere in the world who form or join a team, subject to applicable law. Legal reviewD1
  • Team size (proposal): 2 to 8 human members.D2
  • Age: participants must be of legal age in their country of residence, or have documented consent from a parent or legal guardian. Legal reviewD3
  • Each person may be a member of one team per season.D4
  • Staff of the organizer, judges, and their immediate family may not compete; people linked to partners must disclose that link.D5

3. Team composition: Human + AI + Robotics

Every team must meaningfully combine all three elements. “Meaningfully” means each element does real work in the solution — not a logo on a slide.

  • Human element: the team’s people set goals, make ethical and design decisions, and stay accountable for the outcome.
  • AI element: one or more AI models or tools (e.g. machine learning, computer vision, language or planning models) that contribute analysis, decisions, or optimisation.
  • Robotic element: a system that senses and acts in the physical world with some degree of autonomy or programmed control — for example a drone, rover, robotic arm, sensor-actuator rig, or automated machine.

To keep entry accessible, we propose that low-cost and DIY hardware (e.g. hobby kits, microcontroller-based builds, repurposed equipment) counts fully, and that a high-fidelity simulation of a robot may count where physical hardware is unsafe or out of reach.D6

4. Registration process

  1. Join the early-access list to be notified when registration opens.
  2. Form a team and name a team lead as the main contact.
  3. Choose one challenge track for the season.D7
  4. Submit the registration form: team members, country/region, track, a short plan describing the intended Human–AI–Robot roles, and the public channel(s) for your journey log.
  5. Every member accepts these rules and the code of conduct (and, where needed, guardian consent is provided).
  6. The organizers confirm eligibility; confirmed teams enter the Build & Document phase.

5. Season phases

Calendar dates will be announced later. The season runs in four phases:

Phase 1

Registration

Teams register, pick a track, and are confirmed.

Phase 2

Build & Document

Teams build, test, and keep a public journey log.

Phase 3

Judging

Submissions are verified, audited, and scored; finalists are selected.

Phase 4

World Finals

Finalists present live to global leaders and investors.

6. Submission requirements

By the end of Build & Document, each team submits:

  • Solution description: the problem addressed, the approach, and the role of each Human, AI, and Robotic element.
  • Evidence of outcome: results from a real-world deployment or a documented test (field trial, lab test, or validated simulation), with data, method, and limitations — see the outcome criteria.
  • Public journey log: links to the team’s public updates across the season, including setbacks.D8
  • Technology disclosure: which AI models/tools (including third-party and generative tools) and which robotic hardware or simulators were used, and for what.
  • Safety & ethics statement: risks identified and how they were handled, including any consents obtained.

7. Transparency & documentation

  • Document honestly: show what worked and what failed. Do not overstate results.
  • Label AI-generated or AI-edited media (images, video, voice, text) clearly in public posts.
  • Keep raw evidence (data, logs, unedited footage) available for judges on request.
  • On request, share social-media analytics exports so engagement can be audited.
  • Open-sourcing is encouraged but not required.D9

8. Ethics & safety

  • No harm: solutions and tests must not harm people, animals, or ecosystems. Follow local laws and permits (e.g. for drones, water, or medical items).
  • Physical robot safety: use risk assessments, emergency stops, safe test areas, and supervision for any physical robot, especially around the public.
  • Data privacy: collect only the data you need, protect it, and follow applicable data-protection law.
  • Consent for filming: get consent before filming or posting identifiable people; take extra care with minors and vulnerable people.
  • No deception: no deepfakes, fabricated results, misleading AI content, or fake engagement.

Serious safety or ethics breaches lead to disqualification regardless of score.

9. Intellectual property

Teams keep ownership of their intellectual property. By entering, teams grant the organizers a non-exclusive, royalty-free license to show submission materials (names, descriptions, images, videos, journey-log content) to promote and report on Humanity Games. Teams must have the rights to everything they submit, including third-party and AI-generated material. Legal reviewD10

10. Fair play & disqualification

A team may be warned, penalised, or disqualified for:

  • Buying followers, likes, views, or comments; using bots, engagement pods, or any other engagement manipulation.
  • Plagiarism, or presenting someone else’s work or results as your own.
  • Fabricated or misleading evidence.
  • Safety violations or ignoring safety instructions.
  • Breaching the code of conduct or these rules.

11. Judging & appeals

  • Submissions are checked for eligibility, then scored by an independent panel (see jury selection criteria) using the published scoring criteria and outcome criteria.
  • Judges declare conflicts of interest and step aside where one exists.
  • Teams receive their score breakdown and short written feedback.
  • A team may appeal once, in writing, on grounds of factual error or procedural unfairness, within a set window after results. An appeal reviewer not involved in the original scoring decides; that decision is final.D11

12. World Finals participation

Selected finalists are invited to present at the live World Finals. Format, location, dates, any support for travel, and any awards have not been decided; details will be announced before registration opens. Nothing on this page is a promise of travel funding, prizes, or investment.D12

13. Code of conduct

Be respectful, inclusive, and honest — online and offline, with other teams, judges, organizers, and the public. No harassment, discrimination, hate speech, or intimidation. Credit collaborators and other teams whose ideas you build on. Report concerns to connectwithus@humanitysolutions.net.

14. Liability Legal review

To be drafted with legal counsel before registration opens: liability, insurance, indemnity, and governing-law terms. In the meantime, teams remain responsible for the safety and legality of their own activities.D13

15. Changes to these rules

These rules may change, especially before registration opens. Material changes will be published on this page with a change note; registered teams will be informed. We aim not to make changes during a phase that would unfairly affect teams already competing.

Open decisions

  • D1Which countries or sanctions-related restrictions, if any, limit who may enter.
  • D2Final team size range (proposal: 2–8 human members).
  • D3Minimum age and the process for parent or guardian consent.
  • D4Whether one person may be a member of more than one team in a season.
  • D5Whether employees of partners and sponsors may compete, and on what terms.
  • D6Whether a simulation-only robotic element is allowed, and whether finalists must show physical hardware.
  • D7Whether teams may switch track after registering, and until when.
  • D8Which platforms are accepted for the public journey log, and the minimum update frequency.
  • D9Whether finalists must open-source their code or data.
  • D10Scope and duration of the showcase license granted to the organizers.
  • D11Length of the appeal window and who reviews appeals.
  • D12World Finals format (in-person or hybrid) and whether any travel support is offered.
  • D13Governing law, liability cap, and insurance requirements.
Section 2

Criteria for Choosing Challenge Tracks

How challenge tracks are selected for a season. The four tracks mentioned on the site — Oceanic Cleanup, Future Farming, Urban Housing, and Medical Delivery — are candidates, not final. They will go through the same process as any other proposal.

What makes a good track

  • Global significance & urgency — a real, widely felt problem. The UN Sustainable Development Goals are used as a reference point (Humanity Games is not affiliated with or endorsed by the United Nations).
  • Measurability — outcomes can be measured with clear, verifiable metrics (kg removed, yield gained, time saved, deliveries completed).
  • Genuine need for Human + AI + Robotics — the problem benefits from all three; it is not solvable by software alone or by people alone.
  • Feasibility — a small team can show meaningful progress within one season.
  • Safety & ethics — the track can be run without unacceptable risk to people, ecosystems, or privacy.
  • Accessibility — teams from lower-resource regions can realistically compete (low-cost hardware, open data, simulation options).
  • Public storytelling potential — progress can be shown and understood by a broad audience.
  • Domain partners & data — experts, datasets, or test sites are available to support teams and judging.

Track selection rubric

Reviewers score each candidate track 1 (weak) to 5 (strong) per criterion. Score × weight gives the weighted score. Maximum: 85.

Criterion Score Weight Max
Global significance & urgency1–5×315
Measurability of outcomes1–5×315
Genuine need for Human + AI + Robotics1–5×315
Feasibility within a season for small teams1–5×210
Safety & ethics (a score of 1 excludes the track)1–5×210
Accessibility for lower-resource teams1–5×210
Public storytelling potential1–5×15
Domain partners & data available1–5×15
Total1785

Proposed: tracks need at least 55 of 85 to be shortlisted.D14

Selection process

  1. 1

    Open proposals

    Community members, partners, and domain experts propose tracks with a short brief: problem, metrics, why it needs H-AI-R, safety notes.

  2. 2

    Shortlist

    Organizers score proposals with the rubric above and shortlist those that pass the threshold.

  3. 3

    Expert review

    Independent domain, AI/robotics, and ethics reviewers re-score the shortlist and define track-specific metrics and safety rules.

  4. 4

    Announcement

    Final tracks are published with their briefs, metrics, and safety rules before registration opens.D15

Open decisions

  • D14Minimum rubric score for a track to be shortlisted (proposal: 55 of 85).
  • D15Number of tracks in the inaugural season.
Section 3

Scoring Criteria

Scores are out of 100 points, split evenly: 50% Solution Efficacy and 50% Social Media Impact. Here is what sits inside each half.

50%

Solution Efficacy

50%

Social Media Impact

Scoring table (draft weights)

Sub-criterion What judges look for Points
A · Solution Efficacy — 50 points
A1 Measured impactProgress against the track’s own metrics, compared with a baseline. Outcome criteria15
A2 Evidence quality & verificationSound method, raw data available, results reproducible or independently checked. Evidence levels10
A3 Human–AI–Robot collaborationEach element does real, complementary work; clear human oversight; the combination beats any part alone.12
A4 Scalability & cost-realismRealistic path to wider use; honest cost and maintenance estimates.8
A5 Safety & ethicsRisks identified and handled; privacy and consent respected.5
B · Social Media Impact — 50 points
B1 Normalised reachAudience reached, adjusted for team size, language, and region (see safeguards).10
B2 Genuine engagement qualityMeaningful comments, questions, shares, and follow-up — not raw view counts.12
B3 Storytelling & transparencyA clear, honest journey — including failures and what was learned.12
B4 Educational valueContent that teaches and inspires others to try H-AI-R solutions.10
B5 Community participationInvolving the affected community, volunteers, or other teams; responding to feedback.6
TotalA (50) + B (50)100

Safety is also a gate: a serious safety or ethics breach disqualifies a team regardless of points.

Integrity safeguards

  • Engagement is audited. Finalist candidates share analytics exports; organizers check for anomalies such as sudden spikes, suspicious follower patterns, or engagement pods.
  • Bought or bot engagement disqualifies. Purchased followers, views, likes, or comments, and automated engagement, lead to disqualification.
  • Fair normalisation. Reach is scored relative to comparable teams (team size, language audience size, regional connectivity), so smaller-language and lower-resource teams are not disadvantaged.D16
  • Quality over volume. Most Social Media points (B2–B5) reward quality, honesty, and teaching value, not raw numbers.
  • Balanced judging panel. Each track panel is proposed to include domain experts for that track, AI and robotics practitioners, an ethics/safety specialist, and a science-communication or media specialist — drawn from different regions, with conflicts of interest declared. Judge names will be published once confirmed. Jury selection criteria

Tie-break rules

If teams have equal totals, the tie is broken in this order:

  1. Higher Solution Efficacy subtotal (A).
  2. Higher Human–AI–Robot collaboration score (A3).
  3. Higher evidence quality score (A2).
  4. Higher storytelling & transparency score (B3).
  5. Majority vote of the track’s judging panel.

How finalists are selected (proposal)

  1. Only eligible teams with a complete submission and no integrity issues are ranked.
  2. Teams must reach a minimum Solution Efficacy score so that reach alone cannot carry a team.D17
  3. The top-ranked teams in each track become finalists.D18
  4. The panel may add a limited number of wildcards for outstanding collaboration or teams from under-represented regions.D19
  5. Finalist engagement is audited before invitations are sent.

Open decisions

  • D16Exact method for normalising reach across team size, language, and region.
  • D17Minimum Solution Efficacy score required to become a finalist (e.g. 25 of 50).
  • D18Number of finalists per track.
  • D19Whether wildcard finalists are allowed, and how many.
Section 4

Published criteria: demonstrated real-world outcome toward the chosen challenge

How judges decide whether a team has actually moved its challenge forward. These criteria feed the Solution Efficacy half of the score, mainly A1 Measured impact (15 points) and A2 Evidence quality & verification (10 points).

What counts as an outcome

An outcome is a measurable change toward the track’s challenge — less debris in the water, more food per litre of water, cheaper or faster safe housing, faster and more reliable deliveries. A working prototype, a model, or a plan is valuable, but on its own it is an output, not an outcome.

Output — what you built

  • A drone that can pick up floating plastic
  • An AI model that predicts irrigation needs
  • A deployment plan or simulation

Outcome — what changed

  • Verified kilograms of debris removed from a real site
  • Measured water saved on a real plot vs. a baseline plot
  • Deliveries completed faster than the existing route

Core outcome metrics per track

Before registration opens, each track publishes 2–4 core outcome metrics that every team in that track reports against. Teams may add their own metrics, but must always report the core ones.D20

Illustrative examples only. The metrics below show the kind of measure we mean. They are not final, and they contain no target values or real-world statistics.

Oceanic Cleanup (example)

  • Kg of verified debris removed per operating hour
  • Share of collected material correctly sorted for recycling
  • Wildlife or bycatch incidents (target: none)

Future Farming (example)

  • Yield per litre of water vs. a baseline plot
  • Chemical inputs per area vs. baseline
  • Labour hours per harvest vs. baseline

Urban Housing (example)

  • Cost per housing unit vs. local baseline
  • Build time per unit vs. local baseline
  • Pass rate on structural and safety tests

Medical Delivery (example)

  • Delivery time vs. the existing route
  • Delivery success rate vs. the existing route
  • Temperature compliance for cold-chain items

Baseline and measurement plan (declared up front)

To prevent cherry-picking, each team pre-registers a measurement plan early in the Build & Document phase:

  • Baseline: the current situation you compare against (e.g. the existing route, a control plot, the local cost per unit).
  • Comparison method: how and where you will measure, for how long, and with which tools.
  • Metrics: the track’s core metrics plus any of your own.

The plan is time-stamped and shared with judges. Later changes are allowed but must be logged with a reason; results measured in ways not in the plan count for less.D21

Evidence levels and how they map to points

The strength of your evidence sets the band of points available. Within a band, judges place you by the size of the change against your baseline (A1) and by the quality of data and method (A2), adjusted for scale (see below).

Level What it means A1 Impact /15 A2 Evidence /10
L1Claimed or modelled — estimates, projections, or desk calculations only.0–30–2
L2Lab or simulation tested — controlled test or validated simulation with recorded data.2–62–4
L3Field pilot with data — real-world trial with raw data against the declared baseline.5–104–7
L4Independently verified field result — a partner, expert, or judge has checked the field result.8–137–9
L5Sustained or adopted — a real stakeholder (community, farm, clinic, city) keeps using it.11–159–10

Bands overlap on purpose: a well-documented L2 result can outscore a weak L3 result. Simulation-only entries can reach at most L2.D22

How results are verified

  • Raw data and logs shared with judges (sensor logs, robot telemetry, AI model outputs, measurement sheets).
  • Video with timestamps and location where it is safe and lawful to share; location can be withheld to protect people or sensitive sites.
  • Third-party or partner attestation — a signed confirmation from a site owner, community organisation, or domain partner.
  • Reproducibility check — judges may ask a team to re-run a test or share enough detail for someone else to repeat it.
  • Spot checks and site visits for finalists, in person or by live video.D23
  • AI-content disclosure — any AI-generated or AI-edited evidence (images, video, text, data) must be labelled; undisclosed synthetic evidence is treated as fabrication.

Scale adjustment and honest claims

  • Outcomes are judged relative to the team’s resources and the length of the season. Teams declare their team size, rough budget band, and hardware access at registration.D24
  • Honest small results beat inflated claims. A modest, well-evidenced change scores better than a large claim that cannot be checked.
  • Overclaiming is penalised. Claims that go beyond the evidence lose points in A2.D25
  • Correction process: teams that spot and publicly correct their own overstatement before judging ends are not penalised. Deliberate fabrication leads to disqualification (see Fair play).

Safety & ethics gate

Outcomes achieved by harming people, animals, or ecosystems, by breaking the law, or without the consent of the people affected do not count — however large they are. Serious cases lead to disqualification.

Human–AI–Robot attribution

Teams explain which part of the outcome came from which element: what the humans decided, what the AI contributed, and what the robot physically did. This feeds A3 Human–AI–Robot collaboration (12 points).

Open decisions

  • D20Who sets and approves each track’s 2–4 core outcome metrics.
  • D21Deadline for submitting the measurement plan, and how later amendments are handled.
  • D22Whether work started before the season may count toward evidence levels L4–L5.
  • D23Budget and format (in person or live video) for finalist spot checks and site visits.
  • D24How team resources are declared and used in scale adjustment.
  • D25Size of point deductions for overclaiming.
Section 5

Jury selection criteria

Who judges, how they are chosen, and how they must work. No judges have been appointed yet; names will be published before judging starts.

Panel structure

Each track has its own track panel of 3–5 jurors, and a separate Finals jury scores the World Finals presentations.D26D27

Roles on each track panel (one person may not hold two roles):

Domain expert

Deep knowledge of the track’s challenge (e.g. marine science, agronomy, construction, health logistics).

AI / data expert

Judges the AI contribution, data quality, and measurement method.

Robotics / engineering expert

Judges the physical system, its safety, and its realism.

Ethics, safety, or community representative

Someone from ethics or safety practice, or from a community affected by the challenge.

Storytelling / public engagement expert

Leads the qualitative part of the Social Media Impact half (storytelling, transparency, educational value, community).

Social Media Impact half: the metrics part (normalised reach, engagement authenticity) is audited by a separate engagement-integrity reviewer; the qualitative part is scored by the jury.D28

What we look for in jurors

  • Relevant expertise and track record in their role.
  • Independence from competing teams and from commercial interests in the outcome.
  • Diversity across regions, genders, and disciplines, including representation from the Global South.D29
  • Understanding of low-resource contexts, so frugal solutions are recognised.
  • Availability for calibration, scoring, and deliberation.
  • Commitment to the code of conduct and to confidentiality of team materials. Legal review

Conflicts of interest

  • Every juror declares their interests before appointment and whenever something changes.
  • A juror may not judge a team they advise, fund, employ, work for, or are related to.
  • Sponsors or partners may not be the sole judges of a track they support; their representatives are always a minority on that panel.
  • Recusal: a conflicted juror steps out of scoring and discussion for that team; their score is replaced by the panel average or by a reserve juror.
  • Transparency: the list of jurors and their declared interests is published before judging begins. No names are listed here yet. Legal review

How jurors are selected

  1. 1

    Nominations

    Open nominations from the community and partners, plus invitations from the organizers.D30

  2. 2

    Shortlist

    Candidates are assessed against the criteria above and the role mix each panel needs.

  3. 3

    Independent check

    Someone outside the organizing team checks conflicts of interest and panel balance.D31

  4. 4

    Published roster

    Jurors, roles, and declared interests are published before judging starts.

How jurors work

  • Calibration session: before scoring, each panel scores the same sample submissions and compares results to align on the rubric.
  • Independent scoring first: jurors score alone before any discussion.
  • Per-criterion scores with written rationale for every sub-criterion in the scoring table.
  • Outlier review: where one juror’s score differs strongly from the others, the panel discusses it and the juror may revise with a reason.
  • Blind review where feasible for the Solution Efficacy half: team names and follower counts are removed from written submissions in the first round.D32
  • Audit trail: all scores, rationales, and changes are recorded and kept for appeals.

Compensation & expenses

Whether jurors receive an honorarium or expense reimbursement has not been decided. Any policy will be published and applied equally to all jurors.D33

Removal, replacement & appeals

A juror who breaches the code of conduct, fails to declare a conflict, or cannot complete their duties is removed and replaced by a reserve juror, and affected scores are redone. Teams can challenge results through the appeals process (rule 11).

Open decisions

  • D26Track panel size (3, 4, or 5 jurors).
  • D27Size and make-up of the Finals jury.
  • D28Who acts as the engagement-integrity reviewer (organizer staff, external auditor, or both).
  • D29Diversity commitments for each panel (regions, gender, disciplines, Global South).
  • D30Whether juror nominations use a public form or are by invitation only.
  • D31Who performs the independent conflict-of-interest and balance check on jurors.
  • D32How far blind review is applied in the Solution Efficacy half.
  • D33Juror honorarium and expense-reimbursement policy.

Feedback on this draft is welcome: connectwithus@humanitysolutions.net · Back to top