Log Tempered TF-IDF

From Cohen Courses
Revision as of 02:26, 31 March 2011 by Kdelaros (talk | contribs)
(diff) ← Older revision | Latest revision (diff) | Newer revision → (diff)
Jump to navigationJump to search

Log Tempered TF-IDF is a variant method of calculating the standard information retrieval TF-IDF metric. This metric gives a a weight to how important a word is to a document in a given corpus, and is often used in search engines as part of the scoring / ranking of a document's relevance to a query.

Algorithm / Calculation

Given a document and a corpus, we first calculate the following:

  • Term Frequency: Tf.png
    • A measure of importance of a given term to a document. Frequency of a term for a given document.
  • Inverse Document Frequency: Idf.png
    • A measure of general importance of a term in a corpus

Then the log tempered tf-idf for a word is given by the following:

Lt-tfidf.png

Relevant Papers