Writing and open work
Alongside the peer-reviewed record I write publicly, release open datasets, and build working systems. The through-line is the same as the research: claims about AI should be checkable, and the material needed to check them should be available.
Substack
I publish a newsletter on the real costs of AI: fabricated citations, contaminated data, unmeasured energy, and gaps in governance. It is written for readers who are not specialists but who have to make decisions about these systems.
samaransari.substack.com
Open datasets and code
- PrivHAR-BenchA graduated privacy benchmark for video-based activity recognition. Nine privacy tiers, 1,932 clips, 15 activity classes, pose keypoints, and an evaluation toolkit. Released openly so that privacy-preserving methods can be compared on the same ground.
- FireNetA lightweight fire and smoke detection model, first released in 2019 from supervised undergraduate work. The open dataset remains in active use and the model has been extended as FireNet-v2, FireNet-Tiny, FireNet-Micro and FireNet-Lite.
github.com/SamarAnsariUK
Policy reach
My work on AI slop and data pollution has been cited in WIK Discussion Paper No. 546, in the context of the EU Digital Services Act and disinformation.
Applied projects
- Local redaction of personal data in research transcriptsFunded by the Vice-Chancellor's AI Innovation Fund at Chester. Uses open-weight models on university-controlled hardware to detect and redact personally identifiable information in qualitative interview transcripts, and contributes a method for verifying synthetic data against the data-protection standard of singling out. Preprint in preparation.