โ† All manuals and guides

AlloFlow Dynamic Assessment Studio โ€” Clinician User Guide

Clinical workflow reference ยท reviewed August 20, 2026 ยท operational guidance, not legal advice

What Dynamic Assessment Is

Dynamic Assessment (DA) measures how a student responds to teaching, not just what they already know. Where a traditional test photographs current performance, DA runs a structured cycle โ€” pretest โ†’ mediation โ†’ posttest (with an optional transfer probe) โ€” and watches what changes when you scaffold. The central question is not "Did the student get it right?" but "How much, and how easily, did performance improve when we mediated learning?"

AlloFlow's DA Studio is a technology-supported implementation of this approach. It draws on the established DA tradition โ€” Vygotsky's zone of proximal development (the gap between independent performance and what's possible with support) and Feuerstein's mediated learning experience โ€” operationalizing that tradition with graduated prompt hierarchies and a descriptive growth metric. It does not reproduce or claim equivalence to any specific published instrument (LPAD, ACFS, ARROW, and the like); its methodology is grounded in the tradition but its scoring was designed for this tool and is not normed against any published population.

The Studio is labeled "Clinical tool ยท v1." It supports both an examiner-led path (you control every scaffold) and an AI-mediated path (the model decides which pre-authored scaffold comes next). It is not a substitute for training in dynamic-assessment methodology โ€” it assumes you bring the clinical judgment.


When to Use It โ€” and When Not To

Use DA when you want to:

Do NOT use DA:

Say this plainly in your reports: DA results are clinical observations of learning behavior, not test scores.


The Full Workflow

1. Intake (Referral Context)

The start screen opens with an optional, collapsible "Referral context" panel โ€” six free-text fields, all stored on the device:

The language-background field is load-bearing. Filling it in turns on the access-condition lens described later. Leave it blank for monolingual English students; the lens never appears unless a language background was recorded. Everything you type here surfaces first in your reports and exports. Nothing leaves the device until you choose to send to the Report Writer or invoke AI generation.

You also choose an output language (defaults to your host's leveled-text language), which drives all AI-generated output. Note: the built-in item banks remain in English โ€” to assess in another language, use the Custom Probe builder.

2. Choose a Domain and Build Items

Pick one of four built-in domains, each with banks of 18 items (six easy / six medium / six hard):

Choose a difficulty band (Easy = grades 2โ€“3, Medium = 4โ€“5, Hard = 6โ€“7) and a mediation mode (clinician-led is the default).

Or build a custom probe. The AI builder lets you specify a domain (including free-form "other"), grade band, a required target construct, an optional suspected bottleneck, number of items (3 โ‰ˆ 12 min to 6 โ‰ˆ 30 min), and clinical context (obvious name patterns are soft-stripped before anything reaches the model โ€” a defense-in-depth measure, not a guarantee). It runs a three-pass pipeline โ€” the AI drafts items, critiques its own draft against five quality criteria (catching answer leakage), then refines only the flagged items. You'll see each stage and the critique notes on the review screen. Custom items can be saved to a device-only probe library for reuse.

AI-generated items are not normed or validated. Use them as clinical probes, not standardized measures โ€” and always review before use.

3. Review and Edit

Every generated item is an editable card. You can rewrite the prompt, the correct answer, and acceptable variants, and you can edit each rung of the scaffold ladder individually (regenerating a single rung if needed). The transfer twin (a novel item testing the same construct) is shown read-only. Heuristic warnings flag weak items; the violet AI self-critique block shows the model's own quality verdict, issues, and any refinements (with explicit fallback notices if a critique or refinement pass failed). You can exclude items, regenerate a whole item, and attach supports (see the access section). Then choose Run, Save + Run, or Save only.

4. Run the Session

The session moves through phases automatically:

Pretest (unprompted). One attempt per item, no scaffolds. Record what the student does alone. Scored binary by design โ€” right at level 0 earns full credit, otherwise zero.

Mediation โ€” the 4-level scaffold ladder. This is the heart of DA. When a student struggles, you deliver scaffolds in graduated order, recording the highest level needed:

Each rung can be revealed to yourself first and read aloud to the student. In clinician-led mode you escalate manually and record the level reached โ€” full clinical control. In AI-mediated mode, the model plays examiner: it reads the student's response and decides the next scaffold level, but it always serves the pre-authored ladder text (its role is the decision, never content โ€” no improvised hints). It escalates by exactly one level, never skips, and falls back to a built-in deterministic rule if a call fails (flagged when it does). Note that the leaky-rung correction below applies to the clinician-led path.

Throughout mediation you capture a free-text examiner observation (local-only, never synced), quick-tap observation tags (ten of them: self-corrected, needed wait time, used self-talk, off-task, frustration, asked clarifying question, changed strategy, perseverated, appeared to guess, fluent+automatic), and a leaky-rung flag when a scaffold accidentally gave the answer away.

The ladder card also carries a collapsed "Mediation quality reminders (MLE)" drawer โ€” wait time (5โ€“10 seconds before escalating), intentionality & reciprocity, meaning, and transcendence, drawn from Feuerstein's Mediated Learning Experience criteria. They are reminders that support mediation quality, not a scored fidelity measure. A session-arc stepper (pretest โ†’ mediation โ†’ posttest โ†’ transfer โ†’ results) keeps you oriented in the test-teach-retest cycle throughout.

Posttest (unprompted). Re-test the same items alone. Compare directly to pretest โ€” this is where modifiability shows up.

Transfer (optional). If any item has a transfer twin, the student attempts novel items, same construct, with no scaffolds. This is the test of whether learning generalized rather than was merely memorized.

Answer matching is forgiving for numbers (matches "7" inside "I think 7 apples" but not "17") and uses whole-word containment for text. You can always override with Mark correct / Mark wrong / Skip (or Auto-check + record).

Mis-clicks are recoverable. An โ†ฉ Undo item button re-presents the most recently recorded item with its response, observations, and tags restored โ€” so a slip on "Mark correct" during a live session no longer means discarding everything. It works across phase boundaries, and the results screen has a matching โ†ฉ Reopen last item button that steps back into the session. (For AI-mediated items, the mediation attempts restart fresh on the re-presented item.)

5. Score

Each item scores 5 โˆ’ (prompt level reached): solved alone = 5, after L1 = 4, L2 = 3, L3 = 2, L4 = 1, still wrong after L4 = 0. If you flagged a rung as leaky, a correct response is conservatively credited one level higher โ€” that rung isn't valid evidence of competence. (L4 leaks carry no extra penalty; direct teach states the answer by design. The leaky-rung correction applies to the clinician-led path.)

6. Interpret

The summary screen leads with the Modifiability Index, then layers in transfer, score breakdown, a per-item movement table (each item's arc across pretest โ†’ mediation โ†’ posttest โ†’ transfer, labeled gained / held / not yet / lost), scaffold usage, observation patterns, and โ€” when enabled โ€” the access-condition lens. (Detail below.)

7. Generate Outputs

From the outputs dashboard you can generate a narrative, teacher handoff, family letter, IEP goals, accommodations, and a monitoring plan, then print or export. (Detail below.)


Supports and the Access-Contrast Framework for Multilingual Learners

Supports

Items can carry up to five inline supports (at most one per kind), each anchored to a specific scaffold rung and minted as a clickable resource:

Supports are generated in isolation from any active lesson, standards, or topic โ€” so a support cannot smuggle in outside curriculum content that would confound what the probe is measuring. It tests the construct in the item, not familiarity with whatever lesson happens to be loaded.

Three controlled access contrasts

For multilingual learners (gated on the intake language field), you can run three same-item contrasts. None of them changes the score โ€” each is recorded as a checkbox flag carrying the note "(Access-contrast evidence; does not change the score.)" Each holds the item constant and removes exactly one access demand, so a "flip" (success only in the altered condition) isolates that demand as the likely barrier rather than the underlying skill:

1. Read-aloud (modality). You read the item aloud, then flag "Succeeded only when read aloud." This is direct evidence about whether reading/decoding access โ€” not the reasoning โ€” gates performance.

2. Simpler language (linguistic load). The tool generates a version that reduces only vocabulary and sentence complexity โ€” same numbers, names, quantities, and question, no added hints or steps. Verify the problem is genuinely unchanged before using it, then flag "Succeeded only with simpler language." This isolates academic-language load from the reasoning.

3. Home language (L1). The tool produces a faithful translation at the same complexity. Flag "Succeeded only in home language." This isolates the language of testing from the underlying skill.

The home-language translation is AI-generated. Verify equivalence with a proficient speaker before relying on it โ€” the problem must be unchanged.

Together these triangulate where the barrier sits by elimination. A flip on the read-aloud contrast points at decoding; a flip on simpler-language points at academic-language load; a flip on home-language points at the language of testing. All three are hypothesis-generating observations, never determinations. (The tool deliberately uses no theoretical jargon to label any of this โ€” it describes observations only.)


Interpreting Results

Modifiability Index (MI)

The MI is the proportion of available growth the student realized under mediation:

MI = (posttest โˆ’ pretest) / (maximum possible โˆ’ pretest), bounded from โˆ’1 to 1.

Tiers:

The tier cut-points (0.30 / 0.60) are interpretation conventions of this tool โ€” chosen to structure clinical reasoning, not derived from normative or validation data. Treat the tiers as a reading frame; the underlying pattern (clear gain / partial gain / little change / regression) is the finding.

The MI is descriptive, not normed. With fewer than five items per phase (the default is three), it carries substantial measurement error โ€” read it as a broad direction (clear gain / little change / regression), not a precise number. The summary quantifies this with a sensitivity line: the range the index would span if a single item's scaffold level had been judged one step differently โ€” treat differences smaller than that range as noise. It is a within-student metric: never report it as a standard score, percentile, or classification, and never compare it across students. If you compare against your own saved sessions, the tool will tell you that is a local reference only โ€” not an external or standardized norm โ€” and unstable below ~10 sessions, exploratory below ~30. Across a single student's own sessions the trajectory is descriptive only (and construct-relative โ€” compare points within the same domain), never a growth norm.

Transfer Tier (generalization vs. memorization)

When the transfer phase runs, the tool compares novel-item performance to the posttest:

Recommended report phrasing: "Transfer-probe performance suggests [strong/partial/weak/minimal] generalization to novel surface features."

Learning-Zone Snapshot (ZPD)

The summary also places each administered item into the Vygotskian band the session's own data supports:

A rung you flagged as leaky counts one level higher here (the same conservative correction the score applies), so a leaked L3 success lands in "needs direct teaching" rather than inflating the teachable band. Like everything else on the summary, this is descriptive of this session's items only โ€” not a normed placement.

Access-Condition Lens (hypothesis only)

When a language background is on file, the summary shows a purple "Access-condition lens ยท exploratory" card. It combines two signal types of different epistemic strength:

If language-reducing supports dominate, the lens raises the hypothesis that academic-language access โ€” rather than the underlying skill โ€” limited performance. It is a hypothesis, not a determination. Interpret it only alongside the student's language-proficiency data and home-language context, weigh it against opportunity to learn, and verify any translation with a proficient speaker. It is never a determination of disability or its absence.


Limitations and Responsible Use

Read this section as the floor for every DA report.

In reports: present DA as clinical observations of learning behavior; avoid framing the MI as a standard score, percentile, or classification; do not draw eligibility conclusions from DA alone; and do not frame DA as a "better test" than standardized measures.

For U.S. teams, verify current requirements in 34 CFR ยง300.304, Evaluation procedures and 34 CFR ยง300.300, Parental consent, plus state and district policy. This guide is operational guidance, not legal advice.


Outputs You Can Generate

From the summary screen's outputs dashboard:

Export and backup: a self-contained black-and-white print packet (with the full clinical caveat in its footer), per-session and full-history JSON backup/restore (non-destructive โ€” your local sessions always win), CSV exports for research/instrument development, and a Send to Report Writer hand-off that carries fact chunks, a pre-drafted section, and all generated outputs.


Accessibility and Theming

The Studio renders as a proper dialog (labeled, focus-trapped, Escape closes it from the start screen, and focus returns to the button that opened it). Every focusable control shows a high-contrast focus ring, tables carry proper header semantics, progress bars and charts expose text alternatives, and status changes are announced to screen readers. Destructive actions (discard, delete) confirm through a themed, keyboard-managed dialog rather than the browser's native popup, and Windows High Contrast Mode keeps card structure and focus visible via system colors. The interface follows the app's light / dark / high-contrast theme automatically โ€” including all charts โ€” while printed packets, family letters, and teacher handoffs always print dark-on-white regardless of the on-screen theme. Reduced-motion preferences suppress all animation.


Practical Tips for a Good Session