How to Distinguish Fiction from Nonfiction
Telling fact from fiction isn't always easy on the on the Web. Now researchers have discovered a method that could help automate the process.
kfc 07/22/2010
- 7 Comments

Pick up a piece of text and start reading and it usually becomes clear pretty quickly whether you're reading a nonfictional news story or a fictional novel.
Some clues come from the environment where the stories are found which provide hints, such as the presence of headlines, standfirsts and cross heads.
But even the text alone is revealing. News stories, for example, have very specific structures that give writers little room for creative manoeuvre.
But pinning down these differences in a measurable way that a computer might use to tell them apart is a little more tricky.
Now Joseph Stevanak and Lincoln Carr at the Colorado School of Mines in Golden have come up with a way to do it. They say that the key is to look at the networks that form when you examine how often words appear close together in each type of text.
The type of network they examined creates a graph in which each word in the text forms a vertex. A line connects two vertices if these words appear next to each other in the text. It is possible to explore longer range links by connecting vertices when they appear two or three or four words apart and so on.
Stevanak and Carr say that just two properties of this kind of network can help distinguish fiction from nonfiction stories. The first is the power law that describes the number of links to each vertex in the network. The second is the cluster coefficient which describes how well the vertices are connected to the rest of the network.
Measuring these two quantities alone can identify the type of story with remarkable accuracy. "Our analysis yielded a 73.8±5.15% accuracy for the correct classification of novels and 69.1 ± 1.22% for news stories," say Stevenak and Carr.
This kind of analysis has the potential to improve future generations of text-finding algorithms that they can better classify and hunt down the types of stories that individuals are looking for, and also to identify the communities producing it.
And although it doesn't look like a Google-beater just yet, it has huge potential. If there's one place where the ability to distinguish fact from fiction may turn out to be useful, it is surely on the web.
Ref: arxiv.org/abs/1007.3254: Distinguishing Fact from Fiction: Pattern Recognition in Texts Using Complex Networks



UncleAl
409 Comments
Disinformation's structure
Priority inputs for analysis are the officious utterances of Benjamin "BS" Bernanke, Hank "Hanky-Panky" Paulson, and Tim "Timmy!" Geithner. Next in line would be "Das Capital," "Dianetics," and "The Book of Mormon." Mop up with "Chariots of the Gods" and "Worlds in Collision."
Let's get a neutral call here, then move onto bigger things.
Reply
doanwon
76 Comments
Re: Disinformation's structure
While at it put all the scientific publications containing outlandish theories about the universe through this algorithm. At the same stroke weed out all the self-proclaimed omniscient authority on these theories. Then we can start discussing science seriously.
While it's a good start, I would think that this sort of discriminator would have a hard time distinguishing biographies and autobiographies with historical fiction. With all the knowledge available on the web more attention should be devoted to an intelligent agent embedded into the search engine with basic knowledge that can be used to learn and remember enough to categorize which site is newsworthy (like the major news companies) and which is not, and which is a fictional work vs. a nonfictional one.
Reply
shomas
245 Comments
Re: Disinformation's structure
It needs to be said, differentiating writing styles used in fiction versus nonfiction is not the same as determining truth from lies.
Reply