Difference between revisions of "Class meeting for 10-605 Workflows For Hadoop"

Revision as of 17:29, 18 September 2017

Pig: none required. A nice on-line resource for PIG is the on-line version of the O'Reilly Book Programming Pig.
Optional: Introduction to Information Retrieval, by Christopher D. Manning, Prabhakar Raghavan & Hinrich Schütz, has a fairly self-contained chapter on the vector space model, including Rocchio's method.

Joachims, Thorsten, A Probabilistic Analysis of the Rocchio Algorithm with TFIDF for Text Categorization. Proceedings of International Conference on Machine Learning (ICML), 1997.
Relevance Feedback in Information Retrieval, SMART Retrieval System Experiments in Automatic Document Processing, 1971, Prentice Hall Inc.
Schapire et al, Boosting and Rocchio applied to text filtering, SIGIR 98.

Definition of a similarity join/soft join.
Why inverted indices make TFIDF representations useful for similarity joins
- e.g., whether high-IDF words have shorter or longer indices, and more or less impact in a similarity measure

@@ Line 5: / Line 5: @@
 * First lecture: Slides [http://www.cs.cmu.edu/~wcohen/10-605/workflows-1.pptx in Powerpoint], [http://www.cs.cmu.edu/~wcohen/10-605/workflows-1.pdf in PDF].
 * Second lecture: Slides [http://www.cs.cmu.edu/~wcohen/10-605/workflows-2.pptx in Powerpoint], [http://www.cs.cmu.edu/~wcohen/10-605/workflows-2.pdf in PDF].
+* Third lecture: Slides [http://www.cs.cmu.edu/~wcohen/10-605/workflows-3.pptx in Powerpoint].
-To be updated:
-* Third lecture: Slides [http://www.cs.cmu.edu/~wcohen/10-605/2016/workflow-3.pptx in Powerpoint], [http://www.cs.cmu.edu/~wcohen/10-605/2016/workflow-3.pdf in PDF].
 === Quizzes ===
@@ Line 14: / Line 11: @@
 * [https://qna.cs.cmu.edu/#/pages/view/170 Quiz for first lecture]
 * [https://qna.cs.cmu.edu/#/pages/view/175 Quiz for second lecture]
+* [https://qna.cs.cmu.edu/#/pages/view/178 Quiz for third lecture]
 === Readings ===