Difference between revisions of "N-grams"

Latest revision as of 00:22, 6 November 2012

In the fields of computational linguistics and probability, an n-gram is a contiguous sequence of n items from a given sequence of text or speech. An n-gram could be any combination of letters. However, the items in question can be phonemes, syllables, letters, words or base pairs according to the application. The n-grams typically are collected from a text or speech corpus. An n-gram of size 1 is referred to as a "unigram"; size 2 is a "bigram" (or, less commonly, a "digram"); size 3 is a "trigram". Larger sizes are sometimes referred to by the value of n, e.g., "four-gram", "five-gram", and so on.

@@ Line 1: / Line 1: @@
 In the fields of computational linguistics and probability, an n-gram is a contiguous sequence of n items from a given sequence of text or speech. An n-gram could be any combination of letters. However, the items in question can be phonemes, syllables, letters, words or base pairs according to the application. The n-grams typically are collected from a text or speech corpus.
 An n-gram of size 1 is referred to as a "unigram"; size 2 is a "bigram" (or, less commonly, a "digram"); size 3 is a "trigram". Larger sizes are sometimes referred to by the value of n, e.g., "four-gram", "five-gram", and so on.
+== Relevant Papers ==
+{{#ask: [[UsesMethod::Naive Bayes]]
+| ?AddressesProblem
+| ?UsesDataset
+}}

Difference between revisions of "N-grams"

Latest revision as of 00:22, 6 November 2012

Relevant Papers

Navigation menu

Page actions

Page actions

Personal tools

Navigation

Search

Tools