Empirical evaluations of language-based author identification techniques
International Journal of Speech Language and the Law, Vol. 8, Issue 1, pp. 1–65 (2001)
10.1558/sll.2001.8.1.1
Abstract
Recent Court decisions in the United States call for the empirical testing of language-based author identification techniques. This article shows the results of such testing. The tested hypotheses include: syntactic analysis, syntactically-classified punctuation, sentential complexity, vocabulary richness, readability, content analysis, spelling errors, punctuation errors, word form errors, and grammatical errors. These hypotheses are tested on a set of documents written by four women who are similar in age, educational level, and dialectal background: two of the women are Euro-American, and two are Afro-American. Each hypothesis is tested separately to determine its ability to differentiate documents from different authors and cluster documents from each author. Hypotheses which quantify linguistic features are tested statistically using the chi-square statistic. Discrimination error rates are calculated. Only two hypotheses successfully differentiate and cluster documents: syntactic analysis and syntactically-classified punctuation.
Topics
Field: Computer Science · Subfield: Artificial Intelligence
Keywords
Punctuation,Spelling,Readability,Computer science,Vocabulary,Natural language processing,Linguistics,Syntax,Statistic,Identification (biology)
All Available Versions
- Landing page — International Journal of Speech Language and the Law publishedVersion
- Landing page submittedVersion
Citations by Year
| 2026 | 2025 | 2024 | 2023 | 2022 | 2021 | 2020 | 2019 | 2018 | 2017 |
|---|---|---|---|---|---|---|---|---|---|
| 2 | 2 | 3 | 2 | 7 | 5 | 15 | 28 | 5 | 3 |
Continue on JournalsHub
Put what you just read into practice with these free tools.