Proba · part of the CAQA group Australian vocational education and training

Evidence in education, and how a judgement is proven

Proba is a publication about evidence: how anyone knows that learning happened and that a judgement about a person was sound.

Every certificate, transcript, report card and statement of attainment is a claim about a human being. Somewhere behind that claim sits a body of evidence, a set of rules about what counts, and a person who made a decision. Proba is about that machinery: what evidence actually is in an educational judgement, what makes it hold, and what quietly makes it fail.

Two different claims, two different proofs

There are two sentences that look similar and are not. The first is that the student attended. The second is that the student can do it. The first is a claim about presence, and presence is easy to evidence: a roll, a login record, a timestamp, a signature. The second is a claim about capability, and capability cannot be observed directly at all. It has to be inferred from performance, and inference always carries risk. Confusing the two is the single most common structural error in education, because attendance data is abundant, cheap and tidy, while capability evidence is scarce, expensive and awkward. Systems drift toward the evidence that is easy to collect, and then quietly report it as though it answered the harder question.

An educational judgement is a prediction dressed as a description. When an assessor decides that a person is competent, or that a student has met a learning outcome, the useful meaning is not confined to the moment of assessment. It is that the person will be able to do this again, in a real setting, with a real consequence attached, when nobody is watching for the purpose of marking them. Everything that follows in assessment design exists to make that prediction defensible: the choice of task, the number of observations, the conditions, the record. Evidence is not the paperwork that proves an assessment took place. Evidence is whatever justifies the inference from what was seen to what is claimed.

This is why the word proof needs handling with care. Educational judgement is not proof in the mathematical sense and never will be, because the object being measured is a person operating in conditions that change. What is available is a standard closer to the legal one: a reasoned conclusion, drawn from evidence that a competent independent person could examine and follow to the same place. That is a demanding standard, but it is a reachable one. It requires that the evidence exists, that it was gathered against a stated benchmark, that its limits are understood, and that it survives being looked at later by somebody who was not in the room.

The four rules, explained rather than recited

Australian vocational assessment inherited four rules of evidence: valid, sufficient, authentic and current. They are usually taught as a list to memorise, which is a shame, because each of them exists to rule out a specific and recognisable failure. Read as a set of prohibitions rather than a set of virtues, they become genuinely useful. Validity rules out measuring the wrong thing. Sufficiency rules out concluding from too little. Authenticity rules out crediting the wrong person. Currency rules out relying on a skill that has since decayed or been superseded. Four failures, four guards. An assessment tool that satisfies all four on paper and none of them in practice is common, and the way to detect it is to ask which failure each rule is currently preventing.

Validity is the rule most often lost first, and it fails politely. A written question about a procedure is easy to write, easy to mark and easy to file, but it evidences knowledge of a procedure rather than the ability to perform it. A checklist observation of a simulated task evidences performance under conditions that removed the pressure, the interruptions and the judgement calls that make the real task difficult. Neither is worthless. Both become invalid the moment they are treated as evidence of something they did not test. The test for validity is a direct one: state the claim you intend to make about the person, then ask whether this evidence, on its own, would persuade somebody with an interest in being wrong.

The rules also interact, which the list format hides. Adding more evidence to fix sufficiency can weaken validity if the additional evidence is of the easy sort, because the body of evidence becomes dominated by what was cheap to collect. Tightening authenticity by supervising everything can damage validity where the competency is genuinely performed unsupervised, alone, at a customer site, at three in the morning. Enforcing currency by re-assessing frequently consumes the time that would otherwise buy sufficiency. Assessment design is the practice of losing the least on each of these while remaining honest about what was traded away, and recording that trade so the next person understands why the tool looks the way it does.

Evidence
Anything that supports the inference from what an assessor observed to what the assessor concluded about a person's capability.
Validity
The property of evidence that it actually bears on the thing being claimed, rather than on a convenient proxy for it.
Sufficiency
The property that enough evidence exists, across enough of the required range, to support the judgement without guessing.
Authenticity
The property that the evidence is the work of the person being assessed, and can be shown to be.
Currency
The property that the evidence reflects capability as it stands now, against requirements as they stand now.
Benchmark
The stated standard against which evidence is compared, without which a judgement is only an opinion about quality.

The principles pull against each other

Alongside the rules of evidence sit four principles of assessment: validity, reliability, flexibility and fairness. They are usually presented as a harmonious set, and they are not. They compete, and pretending otherwise produces systems that fail in ways nobody planned. Reliability wants standardisation: the same task, the same conditions, the same marking, so that two assessors reach the same conclusion. Flexibility wants adaptation: different tasks, different timing, different modes, so that assessment fits the learner and the workplace. Every step toward one is a step away from the other. There is no setting of the dial that is correct in general, only settings that are defensible for a particular cohort, qualification and risk profile.

Fairness is not a softening of the other three. It is the requirement that the assessment measures the intended capability and not something incidental to it: reading speed, confidence in English, familiarity with a particular software interface, access to equipment at home, the ability to attend at a fixed hour. An unfair assessment is usually also an invalid one, because it is measuring a second, unstated construct alongside the first. This is the most useful way to argue for reasonable adjustment inside a quality system: it is not a concession made to a learner, it is a correction that restores validity by removing an irrelevant barrier that was contaminating the measurement.

Validity and reliability trade against each other in the most uncomfortable way. The most valid evidence tends to be the messiest: real work, in real conditions, judged holistically by an experienced practitioner. The most reliable evidence tends to be the most reductive: closed questions with a single defensible answer, marked identically by anyone. Push all the way to reliability and you get an assessment that different assessors will score the same and that tells you little about capability. Push all the way to validity and you get rich evidence that two assessors may read differently. Serious systems accept messy evidence and then invest in the human processes, moderation and validation, that keep interpretation of it consistent.

Authenticity is the hardest problem and always was

Authenticity asks a question that no document can settle on its own: is this the work of this person. A submitted assessment is a physical or digital object, and objects do not carry proof of authorship. Signatures attest, they do not demonstrate. Declarations record an intention to be honest. Plagiarism detection compares text to text and finds copying, which is a subset of the problem, not the problem. The uncomfortable truth is that authenticity is never established by the artefact. It is established by the relationship between the artefact and everything else that is known about the person: their prior work, their spoken understanding, their performance when asked to do it again in front of somebody.

The methods that actually work are all variations on that principle. Direct observation of performance answers the question completely for the moment observed. Oral questioning against a submitted piece is powerful because understanding is difficult to borrow: a person who did not produce the work usually cannot defend its choices. Draft histories and staged submissions build a trail in which the finished piece is consistent with what came before. Third party reports from supervisors attest to performance the assessor could not see, and are strongest when they describe specific observed instances rather than confirming general competence. Each method is partial. Used together they narrow the space in which a false attribution can survive.

Online and remote delivery did not create the authenticity problem, it removed the accidental protections that had been masking it. A room full of students under supervision was never a proof of authorship, it was a set of conditions that made substitution inconvenient. When those conditions disappeared, systems that had been relying on inconvenience discovered they had no method at all. Generative tools have applied the same pressure again, and the response that holds up is the same as it always was: assess performance rather than artefacts wherever the competency allows, converse with the person about their work, and treat any single unsupervised submission as one strand of evidence rather than a verdict.

How much evidence is enough, and why more is not better

Sufficiency is a sampling problem, and sampling problems have a shape worth understanding. The claim being made covers a domain: all the situations in which this person will need to perform this skill. The evidence covers a sample of that domain: the situations actually assessed. The judgement is an extrapolation from sample to domain, and the risk of that extrapolation depends less on the number of pieces of evidence than on how well the sample spans the variety inside the domain. One observation is rarely enough, not because one is a small number, but because a single instance cannot show whether the performance was typical, whether it holds under different conditions, or whether it was a fortunate run through a task that usually goes differently.

This is why volume is a poor proxy for sufficiency. Ten written questions about the same narrow aspect of a task provide less coverage than two well chosen pieces of evidence from opposite ends of the required range. Evidence gathered repeatedly under identical conditions confirms consistency in that condition and says nothing about the others. The practical test is to map the required performance, including the variables, contingencies and conditions the standard actually names, and then mark which parts of that map the evidence touches. Blank regions on the map are the sufficiency gap. Duplicated pins in the same region are effort that did not buy any additional confidence.

There is also a real cost to excess. Every additional piece of evidence consumes learner time, assessor attention and storage, and the attention is the scarce resource. Assessors reading a large volume of low value material read it less carefully, and the pieces that mattered get the same brief pass as the padding around them. Oversized evidence requirements also push learners toward compliance behaviour, producing what will satisfy the checklist rather than demonstrating what they can do. Sufficiency achieved through bulk tends to erode validity and reliability at the same time. The discipline is to ask of each requirement what specific gap in the map it closes, and to remove it if the answer is none.

Currency runs in two directions

The usual reading of currency concerns the learner: evidence should reflect what the person can do now, not what they could do at some earlier point. Skills decay, and they decay unevenly. Procedural skills practised daily hold for years. Skills used rarely, particularly ones involving infrequent emergency responses or complex judgement under pressure, decay quickly and often without the person noticing, because confidence outlasts capability. Currency also has a second component that is not about the person at all: the requirement itself may have moved. Legislation changes, equipment changes, safe practice changes. Evidence that was valid against the standard of its time is not evidence against the standard of today, however competently it was gathered.

This matters most in recognition of prior learning, where the entire body of evidence is historical by definition. Recognition done well is not a lighter form of assessment, it is the same assessment conducted against evidence the assessor did not commission. The currency question there is unavoidable and specific: does this evidence describe what the person can do now, and does the practice it describes still match what the standard now requires. Both parts have to be answered. A portfolio of genuinely impressive past work can fail on either, and the honest response is usually supplementary current evidence rather than either wholesale rejection or a generous assumption of continuity.

The second direction is the one systems find easier to neglect: the currency of the assessor. Judging performance against an industry standard requires knowing what that standard looks like in the field today, and that knowledge decays exactly as any other skill does. An assessor who left practice some years ago may still hold the qualification, still know the training product intimately, and still be assessing against a picture of the industry that has quietly gone out of date. This is a validity problem rather than an administrative one, because the benchmark actually applied is the one in the assessor's head. Maintaining it demands real contact with current practice, not a form recording that contact occurred.

If the decision cannot be reconstructed, it is unproven

An assessment decision exists twice: once as an event, and once as a record. The event is gone within minutes. Everything that can be examined afterwards, by a moderator, an auditor, an employer, a court or the learner themselves, lives in the record. This produces a hard practical rule: an assessment judgement that cannot be reconstructed from what was retained is, for every purpose beyond the assessor's memory, unproven. The assessor may have been entirely right. The performance may have been excellent. None of that is available to anyone else. Recognising this changes how records are treated, from an administrative burden that follows assessment to the durable form of the assessment itself.

Reconstruction sets the standard for what a complete record contains. It needs the evidence the learner produced. It needs the benchmark that evidence was compared against, in the version current at the time, because a decision cannot be reviewed against a standard that has since been rewritten. It needs the assessor's judgement expressed with enough specificity to be examined, which means recorded reasoning rather than a tick, particularly at the boundary where a decision could reasonably have gone either way. It needs the conditions: when, where, under what supervision, with what adjustments and why. And it needs the identity and currency of the person who decided.

Version control is the quiet failure point. Tools get revised, sometimes for good reasons, and when the revision overwrites its predecessor the historical record loses the ability to explain itself: past decisions now appear to have been made against a document that did not yet exist. Keeping superseded versions, dated and retrievable, costs almost nothing and preserves the meaning of everything assessed before the change. Retention periods matter for the same reason, and the driver is not filing tidiness but the length of time during which somebody might reasonably need to test the claim. Certification claims outlive the evidence behind them by decades, which is precisely why the evidence has to be findable and legible long after everyone involved has moved on.

Why two good assessors disagree

Two experienced, honest, well trained assessors can look at the same piece of work and reach different conclusions. This is not a scandal and not usually a sign that one of them is incompetent. It happens because judgement requires interpreting a standard written in general language against a performance that is specific, and interpretation varies. Assessors differ in where they set the threshold, in how much weight they give different aspects of a performance, in how they treat an error that was noticed and corrected, and in how much benefit of the doubt they extend. They also differ in what they have seen: an assessor whose recent cohort was unusually strong reads an average performance differently to one whose recent cohort was not.

Some of this variation is reducible and some of it is not. What is reducible comes from ambiguity in the benchmark, from decision rules that were never made explicit, and from assessors having no shared reference point for what the threshold looks like in practice. What is not reducible is the residual judgement that makes the evidence valid in the first place. The response is not to eliminate judgement, which would mean retreating to trivially markable tasks, but to calibrate it, so that the variation left over sits within a range the system can defend and is not correlated with anything it should not be, such as which assessor a learner happened to be allocated.

Moderation and validation do different work and both are needed. Moderation compares judgements: assessors mark the same evidence, surface where they diverge, and argue the divergence out until the threshold is a shared object rather than a private one. Validation examines the instruments and the decisions: whether the tool gathers evidence that meets the four rules, whether the decisions made with it were justified by what was retained, and whether the whole arrangement would satisfy an informed outsider. Both are only useful when the findings change something. Meetings that record agreement and alter no tool, no decision and no practice have consumed the cost of the process and bought none of the benefit. Validation Experts exists for organisations that want that work done properly rather than performed.

Measuring learning, measuring compliance, and the same questions elsewhere

There are two distinct things a quality system can measure, and they are easy to mistake for each other. The first is whether learning occurred and whether judgements about people were sound. The second is whether the required processes were followed and the required artefacts exist. The second is a reasonable proxy for the first, which is why it became the working currency of quality assurance. It is also far easier to check, because process compliance is visible in documents while learning is not. Over time, any system that measures only the proxy will optimise for the proxy, and the optimisation is invisible while it happens, because every indicator continues to look healthy.

The failure mode is specific and recognisable. Tools become longer, because length is defensible. Evidence requirements grow, because more is safer to explain. Records become more complete and less informative, as narrative judgement is replaced by fields that can be audited. Validation meetings become well documented and stop changing anything. Nothing in this sequence involves anyone acting in bad faith. Each step is locally rational. The aggregate result is a system in which the paperwork proves the paperwork, and the original question, whether this person can actually do this, has no owner. The diagnostic is blunt: pick a recent completed learner and ask what in the file would convince a sceptical employer, then read the file honestly. Organisations that want that examined from outside their own habits usually work with CAQA.

None of this is unique to vocational education, though the vocabulary makes it look that way. Schools ask the same questions under different names: moderation of teacher judgement is inter-assessor reliability, work samples and portfolios are evidence sufficiency, and the debate about assessment conditions is a debate about validity against authenticity. Higher education runs the same structure again through academic integrity processes, external examining, constructive alignment between outcomes and tasks, and the persistent tension between assessing rich capability and assessing at scale. The naming differs because the sectors developed their machinery separately. The underlying problem does not differ at all, and practitioners who see the common structure can borrow solutions across boundaries that look impermeable from inside.

Seeing the structure clearly has a practical payoff. It tells you which arguments are actually about evidence and which are about convenience. It tells you when a proposed control will improve a judgement and when it will only make the file thicker. It tells you which parts of an assessment system are load bearing, meaning that removing them would make the judgement unsafe, and which parts exist because somebody once added them and nobody has since asked why. That distinction is the whole practical value of thinking about evidence, and it is available to anyone willing to ask, of every requirement in their system, what failure it prevents.

Every credential is a claim about a person, and the only thing standing behind it is the quality of the evidence and the honesty of the judgement made from it.