What we know
Of 184 responses judged plausible by experts, participants also classified 167 as plausible (90.8%). This measures responses in this test, not deception of real attackers.
Evidence limits
Participants knew the system’s nature; there was no controlled comparison with another honeypot or real-adversary test. The 226 responses are clustered within 12 people and are not independent observations.
Reviewed: full text.
Full-text review · 2026-10-02
Design and scope
Twelve security participants, told they were using a honeypot, assessed 226 shelLM command responses; three experts judged whether those responses were consistent with a Linux shell.
Main finding
Of 184 responses judged plausible by experts, participants also classified 167 as plausible (90.8%). This measures responses in this test, not deception of real attackers.
Evidence limits
Participants knew the system’s nature; there was no controlled comparison with another honeypot or real-adversary test. The 226 responses are clustered within 12 people and are not independent observations.
Locators: PDF pp. 3–5, Sections 4–7, Figure 6, Tables 1–2
The full text was examined and locators were retained. Single-reviewer synthesis without independent extraction or formal quality appraisal.
Sources and provenance
- LLM in the Shell: Generative Honeypots
full-text · 2026-10-02 - Full text examined for Atlas review
full-text · 2026-10-02
Reviewed: 2026-10-02. This record may change when new evidence is found.