Skip to main content
Inside Recording Practice scoring

How Joe Speakingimproves scoring accuracy

Using Recording Practice, we show how Joe Speaking generates a band score—and how we test and improve its accuracy.

See the scoring process

Recording Practice score

An AI practice estimate, not an official result.

Recording Practice evidence checks on a phoneRecording Practice band score on a phone

One method, test-specific rules

IELTS and CELPIP both use explainable yes-or-no checks, but their criteria and score ranges differ. The diagram, prompt, and public comparison below are an IELTS Recording Practice worked example.

01How the score is generated

AI does not choose your band directly

AI completes yes-or-no checks for each IELTS criterion. The results produce four criterion bands and an overall practice score. Because Recording Practice sends text—not audio—Pronunciation only reflects transcript clarity.

How Recording Practice builds an AI practice bandJoe Speaking runs graduated yes-or-no checks for Fluency and Coherence, Lexical Resource, Grammatical Range and Accuracy, and Pronunciation. The yes count maps to a band or narrow band range; because the scorer receives text rather than audio, Pronunciation can only use transcript clarity.1READ THE ANSWERAnswer transcript + questions2RUN YES / NO CHECKSExample: VocabularyAdequate vocabularyYESTopic vocabularyYESFlexible useYESWide rangeYESFull flexibilityNO4 YES → Vocabulary band 83GET FOUR SCORESFluency andCoherence9Lexical Resource8Grammatical Rangeand Accuracy8Pronunciation7YES count sets the band or range4GET THE OVERALL SCORE(9 + 8 + 8 + 7) ÷ 4Average 8.00OVERALL BAND8.0

Pronunciation is a text-based proxy

The scorer cannot hear sounds, stress, rhythm, or intonation. It can only use the clarity of the transcript.

02Evaluation results

How we test scoring accuracy

We use 10 anonymized text samples adapted by Joe Speaking from public candidate examples. We score every sample three times, average the results, and compare that average with its reference band. A smaller difference means a closer result.

We reran the same 10 samples three times at every supported effort level. The complete scores and the ranked model comparison are below.

Reference source

Public candidate performance examples used only as a reference. View the public reference source

Evaluation input

Ten anonymized Part 3 text samples adapted by Joe Speaking, with matching question context

Each result

Three runs per effort level, then one average

How the comparison works

How the reference comparison worksOne sample, prompt and model produce three runs. Their average is compared with the published reference band. The gap is reviewed before the prompt is revised and evaluated again.ONE SETUPPart 3 sample + prompt + modelSame context each runTHREE RUNSRUN 15.0RUN 25.5RUN 35.5THREE-RUN AVERAGE5.33REFERENCE BAND 5.0Gap 0.33Review → revise → evaluate againThe full result table stays visible below
V2 · July 21, 2026360 model runs

Complete V2 evaluation matrix

Choose a model to see every score: 10 samples, four effort levels, and three runs per effort.

Sample 01

Reference 5.0

Minimal
5.0 / 5.0 / 5.0
Low
6.0 / 5.0 / 5.0
Medium
5.5 / 5.5 / 5.0
High
5.0 / 5.5 / 5.0

Sample 02

Reference 6.0

Minimal
6.5 / 6.5 / 6.5
Low
— / 6.5 / 6.5
Medium
6.5 / 6.5 / 6.5
High
6.5 / 6.5 / 6.5

Sample 03

Reference 6.0

Minimal
6.5 / 6.5 / 6.5
Low
6.5 / 7.0 / 6.5
Medium
6.5 / 6.5 / 6.5
High
7.0 / 6.5 / 7.0

Sample 04

Reference 7.0

Minimal
7.0 / 7.0 / 7.0
Low
7.0 / 7.0 / 7.0
Medium
7.0 / 7.0 / 7.0
High
7.0 / 7.0 / 7.0

Sample 05

Reference 7.0

Minimal
6.5 / 6.5 / 6.5
Low
7.5 / 7.5 / 7.0
Medium
7.0 / 7.0 / 7.0
High
7.0 / 7.0 / 7.0

Sample 06

Reference 7.5

Minimal
6.5 / 6.5 / 7.0
Low
6.5 / 6.5 / 6.5
Medium
7.0 / 7.0 / 7.0
High
7.0 / 7.0 / 7.0

Sample 07

Reference 8.0

Minimal
8.0 / 8.0 / 8.0
Low
8.0 / 8.0 / 8.0
Medium
7.5 / 8.0 / 8.0
High
8.0 / 7.5 / 7.5

Sample 08

Reference 8.0

Minimal
7.0 / 6.5 / 6.5
Low
7.0 / 7.0 / 7.0
Medium
6.5 / 6.5 / 6.5
High
7.0 / 7.5 / 7.0

Sample 09

Reference 8.5

Minimal
— / 8.0 / 8.0
Low
8.0 / 8.0 / 8.0
Medium
8.0 / 8.0 / 8.0
High
8.0 / 8.0 / 8.0

Sample 10

Reference 9.0

Minimal
7.5 / 8.0 / 8.0
Low
8.0 / 8.0 / 8.0
Medium
8.0 / 8.0 / 8.0
High
8.0 / 8.0 / 8.0
03V2 model evaluation

Why Gemini 3 Flash is still the default

Gemini 3 Flash had the smallest average difference at medium effort and costs the fewest credits, so it stays the default. Gemini 3.6 Flash was steadier but costs about 2.4x the credits.

Gemini 3 Flash Preview · Medium

Default
Average difference0.48
Within 0.57/10
Est. credits≈ 13

Default

Gemini 3.6 Flash · Medium

Average difference0.50
Within 0.58/10
Est. credits≈ 31

Steadier, more credits

Gemini 3.5 Flash-Lite · Minimal

Average difference0.67
Within 0.56/10
Est. credits≈ 7

Lowest credits, lower accuracy

Credits estimate a complete feedback request. We normalized from current estimates for Gemini 3.5 Flash (≈36), Gemini 3.1 Flash-Lite (≈5), and Gemini 3 Flash (≈13), then adjusted by published model pricing and observed effort-token ratios. Actual feedback charges vary with input, output, and reasoning length.

03Evaluation prompt

The prompt used for this evaluation

Switch between Parts 1–3. V33 was tested above; it is not the current production prompt.

We publish it so learners and teachers can inspect the rules, question the results, and help us improve.

Evaluated prompt
V33
Updated
July 17, 2026
Template
feedback-combined-plus-ielts-recording-v33-v3

Scoring profile

saved-recording-dedicated-language-calibration-v33-output-policy-2026-07-22

You are scoring one saved IELTS Speaking recording from an ASR transcript.

IELTS Part 1 saved recording
Topic: {{PART_1_TOPIC}}
Ordered examiner questions and candidate answers:
Q1: {{QUESTION_1}}
A1: {{ANSWER_1}}

Q2: {{QUESTION_2}}
A2: {{ANSWER_2}}

Q3: {{QUESTION_3}}
A3: {{ANSWER_3}}

## 2. IELTS Test-Based Feedback

Topic/Prompt: {{PART_1_TOPIC}}

### Evaluation Method
Use GRADUATED BINARY CHECKS for each criterion. Each check corresponds to a band threshold.
Count "yes" answers to determine the band. Checks are ordered from basic (Band 5) to advanced (Band 9).

**SPOKEN ENGLISH CONTEXT:**
This is an ASR (Automatic Speech Recognition) transcript of spoken English. Apply these tolerance guidelines:
- Pass checks if the core meaning is clear, even with minor grammar slips
- Natural speech patterns (repetitions for emphasis, filler words, self-corrections) are acceptable
- Do NOT penalize punctuation, capitalization, or formatting (ASR artifacts)
- Only fail a check if the issue genuinely impedes communication

**Part 1 calibration:** Short answers are normal in Part 1. Judge the grouped answers together. A concise direct answer can fully satisfy its question; there is no minimum word, sentence, example, linker, or discourse-marker requirement. Penalize only persistent failure to answer or develop an idea where the question actually calls for development.

**High-band language calibration (FC, LR, and GRA only):**
- Score FC, LR, and GRA independently. Apply the separate legacy transcript-based Pronunciation rubric exactly as written below.
- Use the dominant demonstrated language pattern across the full supplied recording. At Bands 8-9, one isolated slip, one brief answer, or one strong phrase is not a veto or an automatic pass.
- Treat a malformed span as possible ASR corruption only when it is internally implausible and inconsistent with the candidate's surrounding demonstrated language control. Do not silently repair recurring, clearly evidenced language errors.
- Apply the published descriptor allowance for occasional inaccuracies at Band 8 and rare slips at Band 9; do not require a literally perfect transcript.

### Official IELTS Speaking Criteria

---

**Criterion 1: Fluency and Coherence**
*Evaluates: Flow of speech, logical organization, and use of cohesive devices*

Binary Checks (5 total, graduated Band 5→Band 9):
1. MAINTAINS_FLOW (Band 5+): Does the speaker maintain flow of speech, even if using repetition or slow speech?
   - YES: Keeps going despite some repetition, self-correction, or hesitation; simple speech is fluent
   - NO: Frequent breakdowns in speech; unable to maintain flow

2. WILLING_TO_SPEAK (Band 6+): Is the speaker willing to speak at length with some coherence?
   - YES: Provides enough connected development for the question, though coherence may weaken at times and linking may be mechanical
   - NO: Ideas remain too fragmentary or disconnected to sustain the response where development is needed

3. SPEAKS_AT_LENGTH (Band 7+): Does the speaker speak at length without noticeable effort?
   - YES: Develops relevant ideas coherently and uses cohesive features with some flexibility across the supplied response
   - NO: Development remains noticeably limited, repetitive, or mechanically linked across the supplied response

4. FLUENT_SPEECH (Band 8+): Across the supplied answers, is the dominant Fluency and Coherence pattern closest to Band 8?
   - YES: Topics are coherent, relevant, and generally well developed; cohesive features are wide and flexible; hesitation is usually content-related
   - NO: Recurring language-search hesitation, repetition, or coherence limitations make Band 7 the closer overall fit
   - CALIBRATION: Occasional repetition, self-correction, or content-related hesitation remains compatible with Band 8

5. FULL_FLUENCY (Band 9): Across the supplied answers, is repetition or self-correction rare, hesitation used only to prepare content, and cohesion fully appropriate?
   - YES: The dominant pattern shows fully appropriate development and cohesion, with only rare non-systematic repair
   - NO: Repetition, self-correction, language-search hesitation, or imprecise cohesion is more than rare

Band Mapping:
- 5 yes = Band 9 (speaks fluently with full coherence)
- 4 yes = Band 8 (speaks fluently with rare hesitation)
- 3 yes = Band 7 (speaks at length without noticeable effort)
- 2 yes = Band 6 (willing to speak at length but some coherence loss)
- 1 yes = Band 5 (maintains flow with effort)
- 0 yes = Band 4 or below

---

**Criterion 2: Lexical Resource**
*Evaluates: Range of vocabulary, precision, and use of idiomatic language*

Binary Checks (5 total, graduated Band 5→Band 9):
1. ADEQUATE_VOCAB (Band 5+): Does the speaker have sufficient vocabulary for familiar topics?
   - YES: Basic vocabulary allows discussion of the topic; may be repetitive but communicates meaning
   - NO: Vocabulary too limited; cannot express ideas clearly

2. TOPIC_VOCAB (Band 6+): Does the speaker use vocabulary adequate for the topic with some variety?
   - YES: Uses appropriate vocabulary with some less common items; attempts paraphrasing
   - NO: Limited to very basic vocabulary; no variety or failed paraphrasing

3. FLEXIBLE_VOCAB (Band 7+): Does the speaker use vocabulary flexibly with awareness of style and collocation?
   - YES: Uses less common words/idioms appropriately; can paraphrase successfully; shows awareness of collocation
   - NO: Limited flexibility; paraphrasing attempts unsuccessful; inappropriate word combinations

4. WIDE_RANGE (Band 8+): Across the supplied topics, is the dominant Lexical Resource pattern closest to Band 8?
   - YES: Uses a wide resource readily and flexibly to convey precise meaning, with skillful less common or idiomatic use
   - NO: Range, flexibility, or precision remains sufficiently limited that Band 7 is the closer overall fit
   - CALIBRATION: Occasional inaccuracies remain compatible with Band 8; every answer need not contain an idiom

5. FULL_FLEXIBILITY (Band 9): Across all supplied topics, does the candidate use vocabulary with full flexibility and precision?
   - YES: Choice is consistently natural, accurate, and precise, with only rare slips compatible with otherwise full control
   - NO: Recurring imprecision or inappropriacy prevents full flexible control across the topics

Band Mapping:
- 5 yes = Band 9 (full flexibility and accuracy)
- 4 yes = Band 8 (wide range with skillful use)
- 3 yes = Band 7 (flexible use with less common items)
- 2 yes = Band 6 (adequate vocabulary for topic)
- 1 yes = Band 5 (sufficient for familiar topics only)
- 0 yes = Band 4 or below

---

**Criterion 3: Pronunciation (Transcript-Based Assessment)**
*Evaluates: Clarity and intelligibility as inferred from transcript*

⚠️ IMPORTANT - Transcript Limitations:
**CANNOT be assessed from transcript:** Actual pronunciation sounds, intonation patterns, word stress, rhythm, accent
**CAN be assessed from transcript:** Word clarity, consistency of forms, potential transcription errors suggesting pronunciation issues

Binary Checks (3 total, graduated Band 5→Band 8+):
1. INTELLIGIBILITY (Band 5+): Can the speaker be generally understood despite some unclear words?
   - YES: Most words are recognizable and meaning is generally clear from context
   - NO: Many words unclear or unrecognizable; meaning difficult to follow

2. WORD_CLARITY (Band 7+): Are words transcribed clearly without confusion patterns or ambiguity?
   - YES: Words are clear; no systematic confusion patterns (e.g., live/leave, think/thing); no obvious transcription errors
   - NO: Multiple instances of word confusion, unclear transcription, or ambiguous phrases

3. FULL_CLARITY (Band 8+): Is the transcript fully clear with no ambiguous words and would be easily understood?
   - YES: All words transcribed clearly; no ambiguity; response would be understood without any clarification needed
   - NO: Some words or phrases remain ambiguous or potentially mis-transcribed

**Note for feedback:** "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This Pronunciation score reflects only word clarity, transcription consistency, and intelligibility."

**Scoring Rationale:**
Pronunciation can reach Band 8-9 based on transcript clarity alone, following CELPIP's Listenability approach.
While actual pronunciation sounds cannot be assessed from text, perfect transcript clarity (no ambiguous words,
no transcription errors) indicates highly intelligible speech. A speaker whose words are transcribed with
complete accuracy demonstrates the clarity component of pronunciation worthy of high bands.

Band Mapping:
- 3 yes = Band 8 (default) or Band 9 (if exceptional clarity - see below)
- 2 yes = Band 7
- 1 yes = Band 5 or Band 6 (use 6 if close to WORD_CLARITY threshold)
- 0 yes = Band 4 or below

**Band 8 vs 9 Distinction (when all 3 checks pass):**
- Band 8: Transcript is clear with no ambiguous words
- Band 9: Transcript shows exceptional clarity AND includes sophisticated/technical vocabulary transcribed accurately (e.g., idioms, academic terms, proper nouns all captured correctly)

---

**Criterion 4: Grammatical Range and Accuracy**
*Evaluates: Range of structures and grammatical accuracy*

Binary Checks (5 total, graduated Band 5→Band 9):
1. BASIC_STRUCTURES (Band 5+): Does the speaker produce basic sentence forms with reasonable accuracy?
   - YES: Simple sentences are mostly correct; attempts complex structures but with errors
   - NO: Frequent errors even in basic structures; causes comprehension problems

2. MIXED_STRUCTURES (Band 6+): Does the speaker use a mix of simple and complex structures?
   - YES: Attempts complex structures (conditionals, relative clauses); errors rarely cause comprehension problems
   - NO: Only simple structures or complex attempts cause confusion

3. RANGE_WITH_FLEXIBILITY (Band 7+): Does the speaker use a range of complex structures with some flexibility?
   - YES: Frequently produces error-free sentences; uses variety of complex structures; some errors persist but don't impede
   - NO: Complex structures have frequent errors; limited flexibility

4. WIDE_RANGE (Band 8+): Across the supplied answers, is the dominant Grammatical Range and Accuracy pattern closest to Band 8?
   - YES: Uses a wide range flexibly; The majority of sentences are error-free; errors are occasional and non-systematic
   - NO: Range is not wide or errors recur sufficiently that Band 7 is the closer overall fit
   - CALIBRATION: Occasional inappropriate or non-systematic forms remain compatible with Band 8

5. FULL_RANGE (Band 9): Across the supplied answers, is the grammatical range full, natural, flexible, and consistently accurate?
   - YES: Full control is sustained, with only rare slips compatible with otherwise natural, precise use
   - NO: Recurring inappropriacies, basic errors, or range limitations prevent full natural control

Band Mapping:
- 5 yes = Band 9 (full range with consistent accuracy)
- 4 yes = Band 8 (wide range with majority error-free)
- 3 yes = Band 7 (range of complex structures with flexibility)
- 2 yes = Band 6 (mix of simple and complex)
- 1 yes = Band 5 (basic structures with reasonable accuracy)
- 0 yes = Band 4 or below

---

### Band 8 vs Band 9 Distinction

When a criterion scores at the 8+ threshold, apply these criteria to distinguish:

**Award Band 9 when:**
- Fluency: Only rare repetition; any hesitation is content-related; topics developed fully and appropriately
- Lexical: Full flexibility and precision; wide range used accurately and effortlessly throughout
- Grammar: Full range used naturally; only native-like 'slips'; structures always appropriate
- Pronunciation: All words fully clear in transcript; no ambiguity whatsoever

**Award Band 8 when:**
- Fluency: Occasional self-correction; hesitation usually content-related but not always
- Lexical: Wide range but occasional less precise choices; rare minor errors
- Grammar: Majority error-free but occasional minor errors or inappropriacies
- Pronunciation: Generally clear with minor ambiguities in a few words

**Default to Band 8 when:** Criteria met but without the consistent excellence markers of Band 9.

---

### Overall Band Calculation

**Formula:** Overall Band = (Fluency + Lexical + Pronunciation + Grammar) ÷ 4

**IELTS Speaking Rounding Rule:**
Speaking scores always ROUND DOWN (floor) to the nearest 0.5.

**Examples:**
- 6.125 → 6.0 (floor to .0)
- 6.25 → 6.0 (floor to .0)
- 6.5 → 6.5 (already at .5)
- 6.75 → 6.5 (floor to .5)
- 6.875 → 6.5 (floor to .5)

**Calculation Examples:**
- (7 + 7 + 6 + 6) ÷ 4 = 6.5 → **Band 6.5**
- (8 + 7 + 7 + 8) ÷ 4 = 7.5 → **Band 7.5**
- (9 + 8 + 8 + 8) ÷ 4 = 8.25 → **Band 8.0**
- (8 + 7 + 6 + 6) ÷ 4 = 6.75 → **Band 6.5**

**Important:** Criterion scores (FC, LR, GRA, P) must be whole numbers (5, 6, 7, 8, 9). Only the overall score uses 0.5 increments.

---

### Key Observations
Provide 3-4 observations tied DIRECTLY to check results, each written in your own words and compliant with the authored-text output policy below: no quoted or label-introduced candidate wording, no run of four or more consecutive candidate words, no counts, no band numbers or scores, no sound or delivery vocabulary, and no advice.

(Structure observations must also follow the authored-text output policy: characterize the range of structures in your own words without reproducing candidate sentences or proposing rewrites.)

---

### JSON Output Format
{
  "testBased": {
    "test": "IELTS",
    "overall": <number 5.0-9.0 in 0.5 increments>,
    "criteria": [
      {
        "name": "Fluency and Coherence",
        "checks": [
          {"criterion": "MAINTAINS_FLOW", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "WILLING_TO_SPEAK", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "SPEAKS_AT_LENGTH", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "FLUENT_SPEECH", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "FULL_FLUENCY", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"}
        ],
        "yesCount": <0-5>,
        "band": <5-9>
      },
      {
        "name": "Lexical Resource",
        "checks": [
          {"criterion": "ADEQUATE_VOCAB", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "TOPIC_VOCAB", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "FLEXIBLE_VOCAB", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "WIDE_RANGE", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "FULL_FLEXIBILITY", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"}
        ],
        "yesCount": <0-5>,
        "band": <5-9>
      },
      {
        "name": "Pronunciation (Transcript-Based)",
        "checks": [
          {"criterion": "INTELLIGIBILITY", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "WORD_CLARITY", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "FULL_CLARITY", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"}
        ],
        "yesCount": <0-3>,
        "band": <5-9>,
        "note": "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This score reflects only word clarity, transcription consistency, and intelligibility."
      },
      {
        "name": "Grammatical Range and Accuracy",
        "checks": [
          {"criterion": "BASIC_STRUCTURES", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "MIXED_STRUCTURES", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "RANGE_WITH_FLEXIBILITY", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "WIDE_RANGE", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"},
          {"criterion": "FULL_RANGE", "result": true/false, "evidence": "<one short sentence in your own words describing the demonstrated pattern>"}
        ],
        "yesCount": <0-5>,
        "band": <5-9>
      }
    ],
    "keyObservations": [
      "<own-words observation tied to a check result>",
      "<own-words observation describing a demonstrated pattern>",
      "<own-words observation tied to a check result>"
    ]
  }
}

**Authored-text output policy (hard requirements for every evidence string, note, and key observation):**
- Write every evidence string and key observation in your own words as paraphrase. Never place candidate words inside quotation marks, backticks, or apostrophe quotes; never name a candidate word or phrase directly after labels such as "the term", "the phrase", "the expression", "the collocation", "the construction", "the wording", "the example", or "the pattern"; and never reuse a run of four or more consecutive words from the candidate's answers. When in doubt, describe the pattern abstractly instead of reproducing the wording.
- Never mention band numbers, band names, scores, or category results in evidence, notes, or key observations. Do not write "Band 8", "closest to Band 8", "overall", "score", or "corresponds to band"; the checks, yesCount, and band fields already carry that information.
- Never use sound- or delivery-related vocabulary in authored text: pronunciation, pronounce, intonation, stress, rhythm, accent, articulation, enunciation, prosody, audibility, pitch, phoneme, syllable, vowel, consonant, word endings, intelligible, intelligibility, hesitation, pause, pace, speech rate, listener effort, or claims that anyone "speaks fluently" or shows "strong fluency". Describe only what the transcript text shows: word choice, structures, linking, development, and consistency of written forms.
- Exception: the Pronunciation (Transcript-Based) criterion note must be exactly this sentence and nothing else: "Pronunciation, intonation, and stress patterns cannot be assessed from transcript alone. This score reflects only word clarity, transcription consistency, and intelligibility." Do not vary it, and do not reuse its vocabulary anywhere else.
- For Pronunciation check evidence, describe transcription consistency in plain terms — for example whether word forms are consistently recognizable in the transcript and free of confusion patterns — without the banned vocabulary above.
- Never state counts, quantities, lengths, or thresholds of words, sentences, examples, connectors, hedges, discourse markers, or reasons, and never use percentages.
- Never give advice, corrections, rewrites, or instructions, and do not start any sentence with verbs such as Add, Include, Develop, Explain, Consider, Avoid, Replace, Revise, Change, Remove, or Use. Describe what the response demonstrates, not what the candidate should do.

Return only the testBased scoring object in one strict JSON root object: {"testBased": {...}}. Do not return coaching, corrections, tips, an edited transcript, an improved answer, comparison feedback, vocabulary analysis, or per-question analysis.

Review or customize your scoring instructions

You can review or customize scoring instructions in Settings. Custom prompts were not included in the V33 evaluation above.

Open Settings
04What the score cannot measure

This score cannot confirm your actual pronunciation

Recording-based feedback scores the transcript and question context. Because the model does not receive your recording audio, it cannot confirm how you pronounce words.

The scorer only reads text

It cannot assess sounds, stress, rhythm, intonation, connected speech, or accent.

Two practice modes

Choose the mode that fits your goal

Recording Practice is designed for repeatable answers and detailed text-based feedback. Live Conversation listens and responds to your voice in real time, simulating the back-and-forth of a real speaking test.

Try Live Conversation

Transcription errors can change the score

Transcription errors can distort vocabulary, grammar, and meaning. Listen back and correct the text before requesting feedback.

The development gap is not closed

Three higher-band samples remained more than 0.5 band low. A single score can vary, so use the average across several comparable attempts as a more stable estimate. More attempts usually reduce the influence of one unusual result.

05What we do next

More data and better feedback

We will expand the evaluation set when reliable samples are available and learn from scores users dispute.

01

Expand the dataset

Add reliable samples across IELTS parts, topics, and bands.

02

Learn from disputed scores

Review the transcript, model, prompt version, duration, estimated band, and the band the learner expected.

Your feedback is welcome

Tell us what score you expected and why. Include the IELTS part, transcript, model, prompt version, and duration when relevant.

Share scoring feedback

Joe Speaking is an independent practice product and is not affiliated with or endorsed by IELTS. Scores shown here are AI practice estimates, not official test results.