Why two quarters of click rates are not comparable
A click rate belongs to one population, one pretext and one moment. What has to be held constant between rounds, what changes quietly, and how to build an awareness series that supports a comparison.
A click rate is a property of three things at once: the population that received it, the pretext built for them, and the moment it landed. Change any one and the number moves for reasons that have nothing to do with awareness.
So a quarter-on-quarter delta is not a measurement of improvement. It is the sum of what changed in the workforce and what changed in the test, and nothing inside the number separates the two. Two rounds are comparable only if the series was designed to make them so — a decision taken before the first campaign, not a reconciliation attempted after the second.
Pretext difficulty is the dominant variable and the least controlled
Send a round of obvious lures — unknown sender, generic subject, a link resembling nothing the organisation uses — then a round built the way an attacker would build it: a lookalike sender, a system the recipient logs into, a request fitting a process already running that week. The second number will be worse. Nothing about the workforce changed. The test got harder.
It runs the other way too, and that version rarely gets questioned. A careful pretext in round one and a hurried one in round two produce a fall in click rate, and the fall gets reported as progress.
Difficulty is not one dial, but its components are nameable and can be recorded:
- Sender plausibility — an unfamiliar external address, a lookalike domain, or a name the recipient has corresponded with.
- Contextual grounding — whether the message refers to something real: a system in use, a process in flight, a date that means something internally.
- The ask — a bare credential prompt, against a request that fits work already in hand.
- Surface quality and reconnaissance — the ordinary signals a careful reader uses, and how much was known about the recipient beforehand.
Pretexts are built for the client, not drawn from a fixed library, so difficulty cannot be reconstructed from a template name afterwards. Assign a band at design time and record it before results exist, so it cannot be chosen to explain the number it produced. Hold it constant across the rounds being compared and vary content within it, or the series measures recall of one email.
This will be argued about later, so put it in the design note: a programme that raises difficulty as it matures should expect the headline rate to move against it. That is not a regression, and it is only arguable if the band was recorded before the round ran.
The denominator changes quietly
The population measured this quarter is not the one measured last quarter. People join and leave, teams restructure, a business unit moves in or out, and someone redefines a cohort because the old grouping stopped matching the organisation chart.
For a regulated employer the drift is not incidental. RBI's Directions for commercial banks make awareness programmes "mandatory for all new recruits" (Section BB, ¶203, pp. 47–48), so the population is refreshed by obligation. Joiners have the least exposure to the training and to earlier rounds, and their share of the denominator moves on its own.
The NBFC Direction names the fix. ¶36 requires an "up-to-date repository of the training and awareness status of all users" (Section C.12, p. 19). As an evidence artefact that is a compliance obligation; as method it is the precondition for a comparable rate. Without a maintained roll of the population, the denominator is whatever the distribution list held that day.
Fix the cohort definition in advance — function, seniority band, location or entity — and treat a change to it as a break in the series rather than absorbing it. Capture membership at the moment of each send. Publish a reconciliation with every round: count at the last round, joiners, leavers, redefinitions, count now. That line shows whether the headline moved because the cohorts changed or because the cohorts behaved differently. Add the share of each cohort also present last round, because prior exposure moves behaviour on its own.
When the round runs is part of the measurement
The Indian corporate calendar is not flat. Appraisal cycles, financial year end, audit season, quarter close and festival periods each change what an employee expects to receive and how much attention is left for it. A payment pretext arriving in a week of real payment traffic is not the same instrument as the same pretext in a quiet one. A programme that always runs in the same part of the year compares like with like; one whose calendar slips does not, and the slip leaves no mark on the result.
Within the week, time-to-click is only partly a measurement of how quickly people decide. It also measures when mailboxes are being read, so a campaign released at the start of a working day and one released late on a Friday afternoon differ before anyone has exercised judgement.
Two further decisions set the number outright. Clicks accrue for as long as you keep counting, so the counting window has to be fixed and stated with the result. And proximity to another event — a training push, an all-staff communication, a genuine incident, the previous round — produces a number that is real but is not a trend point. Annotate it, do not average it in.
Whether the gateway was allowlisted decides what was measured
Two delivery configurations are defensible and they answer different questions. Allowlist the campaign at the mail gateway and every message reaches a mailbox: what is measured is people, given that this arrived. Leave it unfiltered and what is measured is the whole chain — gateway, filtering, mailbox rules, then the person.
Neither is comparable with the other. Unfiltered, the delivered count is itself an outcome, so a change made by the mail security team between rounds moves the click rate though no employee behaved differently. The improvement is real; it did not come from awareness, and a series that switched configurations credits it there anyway.
The denominator is the same kind of decision. A rate over messages sent, over messages delivered and over messages opened are three numbers with one name. Choose one, record it beside the configuration, and treat a change to either as a new series.
Building a series that survives a comparison
- Assign a difficulty band at design time and compare band against band.
- Fix the cohort definition and publish a reconciliation every round.
- Fix the timing shape — part of the year, release window, counting window — and annotate deviations.
- Record the delivery configuration and the denominator basis as campaign fields.
- Compare like rounds, not adjacent ones. The round that matches on band, cohort and configuration beats the one that merely came before it.
The position to hold from the start: the first round is a baseline, not a verdict. It says nothing about direction, because there is nothing yet for it to have a direction against. Writing that into the first report is easier than defending the number a year later.
An obligation to produce a trend is an obligation to produce a comparable one
The word doing the work across the Indian instruments is periodically. The RBI Cybersecurity, Technology: Risk, Resilience and Assurance Framework Directions, 2026, all issued 31 July 2026, state it for commercial banks at Section BB, ¶202:
"The bank shall evaluate the awareness level of employees periodically."
The same sentence, with the entity named, carries at ¶201 for payments banks, ¶201 for small finance banks and ¶197 for credit information companies. A single measurement does not evaluate a level periodically; it measures one. The obligation is discharged by a series, and a series whose points were produced under different conditions is a set of unrelated measurements wearing a trend line.
The NBFC Direction is more specific. ¶36 requires a "formal mechanism to measure and track the effectiveness of such training through periodic assessments or testing". Track is a word about change over time, and formal mechanism is where the difficulty band, the cohort definition, the timing shape and the delivery configuration live.
SEBI's Cybersecurity and Cyber Resilience Framework requires periodic assessment and names a method by example, at GV.RM Guidelines item 1(e), page 87, standards column GV.RM.S3, applicable to "All REs except small-size, self-certification REs (Mandatory)":
"REs shall periodically assess level of employee cybersecurity awareness, for e.g., through phishing test success rate, etc."
The words for e.g. are the regulator's own. What is named is a rate, assessed periodically.
CERT-In sets the form the result takes. §15.2.2(iii), pp. 55–56 of its Comprehensive Cyber Security Audit Policy Guidelines, CISG-2025-02, Version 1.0, 25 July 2025, requires "anonymized or statistical techniques" when general staff are tested. The population statistic is not a summary of the finding. It is the finding, which makes its construction the whole question.
| Instrument | What the series must carry |
|---|---|
| RBI Directions, 2026 — CB ¶202, PB ¶201, SFB ¶201, CIC ¶197 | Repeated measurement on a stated cadence, so the comparison between rounds is the deliverable |
| RBI Directions, 2026 — NBFC ¶36, p. 19 | Recorded method, and a maintained population roll behind the denominator |
| SEBI CSCRF v1.0, GV.RM Guidelines 1(e), p. 87 | A rate whose basis is stated, so successive rates are one measure |
| CERT-In CISG-2025-02, §15.2.2(iii), pp. 55–56 | Population statistics as the deliverable, their construction documented |
What has to be held is small, and all of it precedes the first campaign: a difficulty band recorded at design time, a fixed cohort definition with a reconciliation line, a fixed timing and counting shape, and a recorded delivery configuration. Four fields on a campaign record, and the difference between a year of numbers and a year of evidence.
About the author
Shalabh Devliyal
Lead — Managed Security Services
Security researcher and penetration tester passionate about making the internet safer. Active CTF player, bug bounty hunter, and hands-on practitioner across web, network, and application security.
Continue reading
All articles →Spear phishing versus bulk simulation: what changes in the test and in the numbers
A spear campaign and a bulk campaign measure different things over different populations. Why their rates cannot share a trend line, and how a six-person cohort is reported when a percentage would identify people.
Scoping a phishing simulation programme in India
What has to be decided before the first send: the population and its cohorts, the scope boundary, the cadence, the channels, the exclusions, the escalation contacts and the written authorisation CERT-In requires.
Measuring the executive population
RBI ¶204 addresses the Board and Senior Management separately, and CERT-In's reporting rule makes a percentage over twelve people an individual result. What an executive exercise reports instead.