Back to projects

SecurityCourse project2026

Keystroke Dynamics

Can AI-generated text be typed out with a human rhythm? A typing simulator built from 99 real typists, tested against the detectors meant to catch it.

99
Real typists in the dataset
0.982
Detector F1, Random Forest
56%
Simulated windows that passed it

Timeline

Winter 2026 term

Role

Research design, ML pipeline, simulator and demo

Team

Solo

Domain

Cybersecurity · Biometric authentication · AI-text detection

Status

Course project

Stack

Pythonscikit-learnSciPypandasFlaskJavaScript

Source

Section 01

Overview

Detectors for AI-written text usually look at the words. A second line of defence looks at how the text was typed: real people have a rhythm of key holds, flights between keys and pauses that a paste or a bot does not. Keystroke dynamics is also used as a behavioural biometric for authentication.

This project asks whether that defence holds. I built a simulator that types any text with timings sampled from real typists, trained detectors to separate human from synthetic typing, and measured how often the simulator gets through. A Flask demo replays the simulated typing live and asks the detectors for a verdict.

The simulator demo
Keystroke Simulator web demo with a text input, speed selector and empty metric cards

Section 02

What Human Typing Looks Like

The data is the KeyRecs free-text dataset: 99 participants, two sessions each, 562,583 key-pair records. Cleaning removed nulls, corrupted rows and pauses over 10 seconds, leaving 559,485.

Three patterns shaped the simulator. Starting a new word is the slowest transition, about 53% slower than typing inside a word (241 ms against 157 ms median). Typists differ by 3.8x in median speed. And 17.2% of transitions are rollovers, where the next key goes down before the last one comes up.

Human timing distributions
Histograms of key hold time, down-down flight time and up-down flight time
Word boundaries slow typing
Box plots of flight time within a word, before a space and after a space
Rollover between keys
Histogram of overlapping and non-overlapping key transitions, and the most common overlapping key pairs

Section 03

The Simulator

For each character, the engine samples a flight time from the distribution fitted to that specific key pair, then adjusts it for context (word start, mid-word, word end), the chosen speed profile and slow fatigue drift. It blends each new flight with the previous one so rhythm has momentum, and adds thinking pauses after commas and full stops. Hold times are sampled per key.

The output is a full keystroke stream with key-down and key-up times, which the demo replays in real time while charting every flight.

A finished simulation
Completed simulation with typing metrics and a chart of flight time per keystroke
Keystroke simulator demo on a phone
The demo on a phone

Section 04

Detection

Detectors see 19 features computed over windows of 20 keystrokes: flight and hold statistics, their variability, rollover ratio and hold-to-flight ratio. I trained Random Forest, Gradient Boosting and AdaBoost on 26,430 human windows.

The first version scored almost perfectly, which was the warning sign: its synthetic negatives were too easy to separate. I rebuilt the negatives from statistical mimics, independently resampled features and noise-perturbed human windows. On that harder task Random Forest reached F1 0.982 and AUC 0.999, Gradient Boosting 0.901 and AdaBoost 0.745. K-Means on the participants found four typing archetypes, from steady to fast-and-overlapping.

Three detectors compared
ROC curves, classification metrics and top feature importances for three detectors
Four typing archetypes
Elbow plot, PCA scatter and sizes of four typing archetypes

Section 05

Results

Against the strongest detector the simulator is close to a coin flip: Random Forest called 56.3% of simulated windows human. Gradient Boosting caught most of them, passing only 23.5%. AdaBoost passed 82.6%, but it also passes 100% of naive synthetic typing, so that number says more about AdaBoost than about the simulator. Naive synthetic typing was caught almost every time by the two strong detectors (0% and 0.5% passed).

Asking the detectors
Demo evaluation panel with Random Forest leaning human and Gradient Boosting calling the run synthetic
Simulation against each detector
Share of simulated and naive windows classified human by each detector

Section 06

What Gave It Away

My report attributed the misses to rollover, which the simulator does not model explicitly. Going back to the outputs, rollover was actually close to human, because short sampled flights overlap long holds on their own. The clearer tell is hold time: some fitted distributions were degenerate, so about 18% of simulated holds sit at the 20 ms floor, and hold features are four of the Random Forest's top seven.

The fixes are concrete: refit holds with a fixed location or sample them empirically, add typos and corrections, and train detectors on keystroke-level synthetic streams instead of synthetic feature vectors.

Real against simulated timing
Real and simulated flight and hold time histograms, with a Q-Q plot of flight times