Bag of Words

Others Dataset

Homepage

http://archive.ics.uci.edu/ml/datasets/Bag+of+Words

Description

For each text collection, D is the number of documents, W is the number of words in the vocabulary, and N is the total number of words in the collection (below, NNZ is the number of nonzero counts in the bag-of-words). After tokenization and removal of stopwords, the vocabulary of unique words was truncated by only keeping words that occurred more than ten times. Individual document names (i.e. a identifier for each docID) are not provided for copyright reasons. These data sets have no class labels, and for copyright reasons no filenames or other document-level metadata. These data sets are ideal for clustering and topic modeling experiments. For each text collection we provide docword.*.txt (the bag of words file in sparse format) and vocab.*.txt (the vocab file). Enron Emails: orig source: www.cs.cmu.edu/~enron D=39861 W=28102 N=6,400,000 (approx) NIPS full papers: orig source: books.nips.cc D=1500 W=12419 N=1,900,000 (approx) KOS blog entries: orig source: dailykos.com D=3430 W=6906 N=467714 NYTimes news articles: orig source: ldc.upenn.edu D=300000 W=102660 N=100,000,000 (approx) PubMed abstracts: orig source: www.pubmed.gov D=8200000 W=141043 N=730,000,000 (approx)

Discussion

Related datasets

NIPS Conference Papers 1987-2…

The dataset is in the form of a 11463 x 5812 matrix of word counts, containing 11463 words and 5811 NIPS conference papers (the first column contains the …

clustering, text

Computer Science

3D Road Network (North Jutlan…

This dataset was constructed by adding elevation information to a 2D road network in North Jutland, Denmark (covering a region of 185 x 135 km^2). Elevati…

clustering, regression, sequential, text

Computer Science

Dresses_Attribute_Sales

Style, Price, Rating, Size, Season, NeckLine, SleeveLength, waiseline, Material, FabricType, Decoration, Pattern, Type, Recommendation are Attributes in d…

classification, clustering, text

Computer Science

KEGG Metabolic Reaction Netwo…

KEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. In …

classification, clustering, multivariate, regression, text, univariate

Life Sciences

Reuter_50_50

The dataset is the subset of RCV1. These corpus has already been used in author identification experiments. In the top 50 authors (with respect to total s…

classification, clustering, domain-theory, multivariate, text

Computer Science

KEGG Metabolic Relation Netwo…

KEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. In …

classification, clustering, multivariate, regression, text, univariate

Life Sciences

TTC-3600: Benchmark dataset f…

The dataset consists of a total of 3600 documents including 600 news/texts from six categories economy, culture-arts, health, politics, sports and techno…

classification, clustering, text

Computer Science

NYSK

Documents are first obtained via a Web search using AMIEI: an integrated platform for delivering enterprise intelligence, developed by AMI Software ([Web …

clustering, multivariate, sequential, text

Social Sciences

Miskolc IIS Hybrid IPS

The measurements were created to ease the development, comparison and evaluation of fingerprinting based hybrid indoor positioning methods. The measuremen…

causal-discovery, classification, clustering, text

Computer Science

YouTube Multiview Video Games…

Please see the README for the details on the data organization, and so on.

classification, clustering, multivariate, text

Computer Science

SMS Spam Collection

This corpus has been collected from free or free for research sources at the Internet: -> A collection of 425 SMS spam messages was manually extracted fr…

classification, clustering, domain-theory, multivariate, text

Computer Science

Airport MotionSeg

The Airport MotionSeg dataset contains 12 sequences of videos of an aiprort scenario with small and large moving objects and various speeds. It is challen…

airport, camera, clustering, motion, segmentation, video, zoom

Vision

ChokePoint Dataset

We collected a video dataset, termed ChokePoint, designed for experiments in person identification/verification under real-world surveillance conditions u…

clustering, detection, face, human, identification, multiview, pedestrian, real, recognition, sequence, surveillance, world

Vision

Activities of Daily Living (A…

This dataset comprises information regarding the ADLs performed by two users on a daily basis in their own homes. This dataset is composed by two instanc…

classification, clustering, multivariate, sequential, time-series

Computer Science

Sales_Transactions_Dataset_We…

52 columns for 52 weeks; normalised values of provided too.

clustering, multivariate, time-series

Others

wiki4HE

Ongoing research on university faculty perceptions and practices of using Wikipedia as a teaching resource. Based on a Technology Acceptance Model, the re…

causal-discovery, clustering, multivariate, regression

Social Sciences

Anuran Calls (MFCCs)

This dataset was used in several classifications tasks related to the challenge of anuran species recognition through their calls. It is a multilabel data…

classification, clustering, multivariate

Life Sciences

Sentiment Labelled Sentences

This dataset was created for the Paper 'From Group to Individual Labels using Deep Features', Kotzias et. al,. KDD 2015 Please cite the paper if you want …

classification, text

Others

Gas Sensor Array Drift Datase…

This data set contains 13,910 measurements from 16 chemical sensors exposed to 6 gases at different concentration levels. This dataset is an extension of …

causa, classification, clustering, multivariate, regression, time-series

Computer Science

Twenty Newsgroups

N/A

text

Others

YACCLAB dataset

The YACCLAB dataset includes both synthetic and real binary images and is suitable for a wide range of applications, ranging from document processing to s…

binary, fingerprints, labeling, medical, natural, randomnoise, text, videosurveillance

Vision

User Knowledge Modeling

-- The users' knowledge class were classified by the authors using intuitive knowledge classifier (a hybrid ML technique of k-NN and meta-heuristic…

classification, clustering, multivariate

Computer Science

Grammatical Facial Expressions

The automated analysis of facial expressions has been widely used in different research areas, such as biometrics or emotional analysis. Special import…

classification, clustering, multivariate, sequential

Computer Science

YouTube Spam Collection

The table below lists the datasets, the YouTube video ID, the amount of samples in each class and the total number of samples per dataset. Dataset --- Yo…

classification, text

Computer Science

Burst Header Packet (BHP) flo…

For Further information about the variables see the file in the data folder.

classification, text

Computer Science

Daily and Sports Activities

Brief Description of the Dataset: --------------------------------- Each of the 19 activities is performed by eight subjects (4 female, 4 male, between th…

classification, clustering, multivariate, time-series

Computer Science

Plants

The data is in the transactional form. It contains the Latin names (species or genus) and state abbreviations.

clustering, multivariate

Life Sciences

seeds

The examined group comprised kernels belonging to three different varieties of wheat: Kama, Rosa and Canadian, 70 elements each, randomly selected for the…

classification, clustering, multivariate

Life Sciences

Street View House Number (SVH…

SVHN is a real-world image dataset for developing machine learning and object recognition algorithms with minimal requirement on data preprocessing and fo…

classification, detection, number, real, recognition, streetside, streetview, text, urban, world

Vision

Twitter Data set for Arabic S…

--- By using a tweet crawler, we collect 2000 labelled tweets (1000 positive tweets and 1000 negative ones) on various topics such as: politics an…

classification, text

Social Sciences

MoCap Hand Postures

A Vicon motion capture camera system was used to record 12 users performing 5 hand postures with markers attached to a left-handed glove. A rigid patter…

classification, clustering, multivariate

Computer Science

Reuters Transcribed Subset

Data Characteristics: -------------------- This data was created by selecting 20 files each from the 10 largest classes in the Reuters-21578 collection …

classification, text

Business

Yahoo Flickr Creative Commons…

Yahoo Flickr Creative Commons 100M (YFCC100M) dataset contains a list of photos and videos. This list is compiled from data available on Yahoo! Flickr. Al…

3d, clustering, community, detection, flickr, image, internet, landmark, recognition, reconstruction, social

Vision

StoneFlakes

Background information: The data set concerns the earliest history of mankind. Prehistoric men created the desired shape of a stone tool by striking on a …

causal-discovery, classification, clustering, multivariate

Others

Reuters-21578 Text Categoriza…

From the original readme file (please consult it for more information): ------------------------- The documents in the Reuters-21578 collection appeared o…

classification, text

Others

Tennis Major Tournament Match…

N/A

classification, clustering, multivariate, regression

Others

User Identification From Walk…

The dataset collects data from an Android smartphone positioned in the chest pocket. Accelerometer Data are collected from 22 participants walking in the …

classification, clustering, sequential, time-series, univariate

Others

Online Retail

This is a transnational data set which contains all the transactions occurring between 01/12/2010 and 09/12/2011 for a UK-based and registered non-store o…

classification, clustering, multivariate, sequential, time-series

Business

Individual household electric…

This archive contains 2075259 measurements gathered between December 2006 and November 2010 (47 months). Notes: 1.(global_active_power*1000/60 - sub_mete…

clustering, multivariate, regression, time-series

Physical Systems

Text and Vision (TVGraz) Data…

The Text and Vision (TVGraz) dataset is an annotated multi-modal dataset which currently contains 10 visual object categories, 4030 images and associated …

appearance, classification, evaluation, text

Vision

Water Treatment Plant

This dataset comes from the daily measures of sensors in a urban waste water treatment plant. The objective is to classify the operational state of the pl…

clustering, multivariate

Physical Systems

OpinRank Review Dataset

Car Reviews ------------ -Full reviews of cars for model-years 2007, 2008, and 2009 -There are about 140-250 cars for each model year -Extracted fields in…

text

Computer Science

Northix

Northix is designed to be a schema matching benchmark problem for data integration of two entity relationship databases. Northix is the resulting schema m…

classification, multivariate, text, univariate

Computer Science

Educational Process Mining (E…

The experiments have been carried out with a group of 115 students of first-year, undergraduate Engineering major of the University of Genoa. We carried…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

AAAI 2013 Accepted Papers

CSV format where each row is a paper and each column is an attribute.

clustering, multivariate

Computer Science

Motion Capture Hand Postures

A Vicon motion capture camera system was used to record 12 users performing 5 hand postures with markers attached to a left-handed glove. A rigid patter…

classification, clustering, multivariate

Computer Science

ElectricityLoadDiagrams201120…

Data set has no missing values. Values are in kW of each 15 min. To convert values in kWh values must be divided by 4. Each column represent one client. S…

clustering, regression, time-series

Computer Science

microblogPCU

Our dataset is used by us to explore spammers in microblog and you can access our demo system at [Web Link]Please add :8080 after the domain name as port.…

causal-discovery, classification, multivariate, sequential, text, univariate

Computer Science

Sponge

These are atlantic-mediterranean marine sponges that belong to O.Hadromerida (Demospongiae.Porifera).

clustering, multivariate

Life Sciences

DBWorld e-mails

I collected 64 e-mails from DBWorld newsletter and I used them to train different algorithms in order to classify between 'announces of conferences' and '…

classification, text

Computer Science

Diabetes 130-US hospitals for…

The dataset represents 10 years (1999-2008) of clinical care at 130 US hospitals and integrated delivery networks. It includes over 50 features representi…

classification, clustering, multivariate

Life Sciences

Synthetic Control Chart Time …

This dataset contains 600 examples of control charts synthetically generated by the process in Alcock and Manolopoulos (1999). There are six different cla…

classification, clustering, time-series

Others

NSF Research Award Abstracts …

The abstracts, one per file, were furnished by the NSF (National Science Foundation). A sample abstract is shown in the next section. The bag-of-word dat…

text

Others

Dataset for ADL Recognition w…

The Dataset for ADL Recognition with Wrist-worn Accelerometer is a public collection of labelled accelerometer data recordings to be used for the creation…

classification, clustering, multivariate, time-series

Computer Science

Amazon Commerce reviews set

dataset are derived from the customers reviews in Amazon Commerce Website for authorship identification. Most previous studies conducted the identifi…

classification, domain-theory, multivariate, text

Physical Systems

UJIIndoorLoc-Mag

Indoor localization is a key topic for mobile computing. However, it is still very difficult for the mobile sensing community to compare state-of-art Indo…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

Parkinson Disease Spiral Draw…

The PD and control handwriting database consists of 62 PWP (People with parkinson) and 15 healthy individuals who appealed at the Department of Neurology …

classification, clustering, multivariate, regression

Computer Science

Geo-Magnetic field and WLAN d…

Indoor localisation is a key topic for the Ambient Intelligence (AmI) research community. In this scenarios, recent advancements in wearable technologie…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

Dow Jones Index

In predicting stock prices you collect data over some period of time - day, week, month, etc. But you cannot take advantage of data from a time period unt…

classification, clustering, time-series

Business

KDC-4007 dataset Collection

The most important feature of this dataset is its simplicity to use and its being well-documented, which can be widely used in various studies of text ana…

classification, multivariate, regression, text

Computer Science

Eco-hotel

The CSV holds one cell per review, except the first row, which is the header. All the 401 reviews were collected between January and August of 2015.

text

Business

Libras Movement

The dataset (movement_libras) contains 15 classes of 24 instances each, where each class references to a hand movement type in LIBRAS. In the video pre-p…

classification, clustering, multivariate, sequential

Others

Farm Ads

This data was collected from text ads found on twelve websites that deal with various farm animal related topics. Information from the ad creative and th…

classification, text

Business

Activity Recognition from Sin…

--- The dataset collects data from a wearable accelerometer mounted on the chest --- Sampling frequency of the accelerometer: 52 Hz --- Accelerom…

classification, clustering, sequential, time-series, univariate

Others

Idiap/ETHZ Faces and Poses

Idiap/ETHZ Faces and Poses Dataset dataset by L. Jie, B. Caputo and V. Ferrari contains 1703 image-caption pairs. [author] Captions contain the names of s…

face detection, object pose, pedestrian, text

Vision

News Aggregator

News are grouped into clusters that represent pages discussing the same news story. The dataset includes also references to web pages that, at the access…

classification, clustering, multivariate

Others

Perfume Data

The data set gathered when we were working at project for Bahrain university between 2002 and 2003.

classification, clustering, domain-theory, univariate

Computer Science

TV News Channel Commercial De…

Automatic identification of commercial blocks in news videos finds a lot of applications in the domain of television broadcast analysis and monitoring. Co…

classification, clustering, multivariate

Computer Science

Tamilnadu Electricity Board H…

Collect the real time readings for residential,commercial,industrial,agriculure,to find the accuracy consumption in Tamil Nadu Around Thanajvur

classification, clustering, multivariate, regression

Life Sciences

Taxi Service Trajectory - Pre…

For complete information see the official challenge page: [Web Link]

causal-discovery, clustering, domain-theory, multivariate, sequential, time-series

Computer Science

Badges

Part of the problem in using an automated program to discover the unknown target function is to decide how to encode names such that the program can be us…

classification, text, univariate

Others

DrivFace

The DrivFace database contains images sequences of subjects while driving in real scenarios. It is composed of 606 samples of 640480 pixels each, acquired…

classification, clustering, multivariate, regression

Computer Science

Epileptic Seizure Recognition

Please find the original data at '[Web Link]'

classification, clustering, multivariate, time-series

Life Sciences

CNAE-9

This is a data set containing 1080 documents of free text business descriptions of Brazilian companies categorized into a subset of 9 categories cataloged…

classification, multivariate, text

Business

Human Activity Recognition Us…

The experiments have been carried out with a group of 30 volunteers within an age bracket of 19-48 years. Each person performed six activities (WALKING, W…

classification, clustering, multivariate, time-series

Computer Science

Open University Learning Anal…

Open University Learning Analytics Dataset (OULAD) contains data about courses, students and their interactions with Virtual Learning Environment (VLE) fo…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

All I Have Seen (AIHS)

The All I Have Seen (AIHS) dataset is created to study the properties of total visual input in humans, for around two weeks Nebojsa Jojic wore a camera ca…

3d, clustering, indoor, outdoor, scene, similarity, study, summary, user, video

Vision

SML2010

The dataset could contain missing values. The data was sampled every minute, computing and uploading it smoothed with 15 minute means. The header of the d…

multivariate, regression, sequential, text, time-series

Computer Science

HTRU2

HTRU2 is a data set which describes a sample of pulsar candidates collected during the High Time Resolution Universe Survey (South) [1]. Pulsars are a r…

classification, clustering, multivariate

Physical Systems

Folio

- The leaves were placed on a white background and then photographed. - The pictures were taken in broad daylight to ensure optimum light intensity.

classification, clustering, multivariate

Others

AAAI 2014 Accepted Papers

CSV format where each row is a paper and each column an attribute.

clustering, multivariate

Computer Science

Opinosis Opinion Review

This dataset contains sentences extracted from user reviews on a given topic. Example topics are performance of Toyota Camry and sound quality of ipod nan…

text

Computer Science

YouTube Comedy Slam Preferenc…

YouTube Comedy Slam ([Web Link]) is a video discovery experiment running on YouTube's version of labs (called TestTube) for a few months in 2011 and 2012.…

classification, text

Computer Science

Wholesale customers

Provide all relevant information about your data set.

classification, clustering, multivariate

Business

Turkiye Student Evaluation

N/A

classification, clustering, multivariate

Others

Legal Case Reports

This dataset contains Australian legal cases from the Federal Court of Australia (FCA). The cases were downloaded from AustLII ([Web Link]). We included a…

classification, text

Others

Mice Protein Expression

The data set consists of the expression levels of 77 proteins/protein modifications that produced detectable signals in the nuclear fraction of cortex. Th…

classification, clustering, multivariate

Life Sciences

US Census Data (1990)

The data was collected as part of the 1990 census. There are 68 categorical attributes. This data set was derived from the USCensus1990raw data set. The…

clustering, multivariate

Social Sciences

FMA: A Dataset For Music Anal…

* Audio track (encoded as mp3) of each of the 106,574 tracks. It is on average 10 millions samples per track.* Nine audio features (consisting of 518 attr…

classification, clustering, multivariate, time-series

Computer Science

Gesture Phase Segmentation

The dataset is composed by features extracted from 7 videos with people gesticulating, aiming at studying Gesture Phase Segmentation. Each video is repres…

classification, clustering, multivariate, sequential, time-series

Others

Character Trajectories

The characters here were used for a PhD study on primitive extraction using HMM based models. The data consists of 2858 character samples, contained in th…

classification, clustering, time-series

Computer Science

Heterogeneity Activity Recogn…

The Heterogeneity Dataset for Human Activity Recognition from Smartphone and Smartwatch sensors consists of two datasets devised to investigate sensor het…

classification, clustering, multivariate, time-series

Computer Science

Syskill and Webert Web Page R…

The HTML source of a web page is given. Users looked at each web page and inidated on a 3 point scale (hot medium cold) 50-100 pages per domain. However, …

classification, multivariate, text

Computer Science

Sentence Classification

Please see the README file that accompanies the data.

classification, text

Others

UNIX User Data

This file contains 9 sets of sanitized user data drawn from the command histories of 8 UNIX computer users at Purdue over the course of up to 2 years (USE…

sequential, text

Computer Science

gene expression cancer RNA-Seq

Samples (instances) are stored row-wise. Variables (attributes) of each sample are RNA-Seq gene expression levels measured by illumina HiSeq platform.

classification, clustering, multivariate

Life Sciences

Amazon Access Samples

This is a sparse data set, less than 10% of the attributes are used for each sample. The link is to a '*.tgz' file which contains two files: [amzn-anon-ac…

causal-discovery, clustering, domain-theory, regression, time-series

Business

Bag of Words

Homepage

Description

Tags

Discussion

Related datasets