The CSV holds one cell per review, except the first row, which is the header. All the 401 reviews were collected between January and August of 2015.
For each text collection, D is the number of documents, W is the number of words in the vocabulary, and N is the total number of words in the collection...
text, clusteringThe abstracts, one per file, were furnished by the NSF (National Science Foundation). A sample abstract is shown in the next section. The bag-of-word d...
textThis corpus has been collected from free or free for research sources at the Internet: -> A collection of 425 SMS spam messages was manually extracted ...
text, classification, clustering, multivariate, domain-theoryThe dataset is in the form of a 11463 x 5812 matrix of word counts, containing 11463 words and 5811 NIPS conference papers (the first column contains th...
text, clusteringFor Further information about the variables see the file in the data folder.
text, classificationThe YACCLAB dataset includes both synthetic and real binary images and is suitable for a wide range of applications, ranging from document processing to...
fingerprints, videosurveillance, text, binary, medical, natural, labeling, randomnoiseThe dataset is the subset of RCV1. These corpus has already been used in author identification experiments. In the top 50 authors (with respect to total...
text, classification, clustering, multivariate, domain-theoryThe dataset could contain missing values. The data was sampled every minute, computing and uploading it smoothed with 15 minute means. The header of the...
text, sequential, regression, time-series, multivariateOur dataset is used by us to explore spammers in microblog and you can access our demo system at [Web Link]Please add :8080 after the domain name as por...
univariate, classification, sequential, causal-discovery, multivariate, textData Characteristics: -------------------- This data was created by selecting 20 files each from the 10 largest classes in the Reuters-21578 collection...
text, classificationThis dataset contains Australian legal cases from the Federal Court of Australia (FCA). The cases were downloaded from AustLII ([Web Link]). We included...
text, classificationKEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. I...
univariate, classification, clustering, regression, multivariate, textThe Text and Vision (TVGraz) dataset is an annotated multi-modal dataset which currently contains 10 visual object categories, 4030 images and associate...
text, evaluation, appearance, classificationNorthix is designed to be a schema matching benchmark problem for data integration of two entity relationship databases. Northix is the resulting schema...
multivariate, text, univariate, classificationThe measurements were created to ease the development, comparison and evaluation of fingerprinting based hybrid indoor positioning methods. The measurem...
text, classification, clustering, causal-discoveryIdiap/ETHZ Faces and Poses Dataset dataset by L. Jie, B. Caputo and V. Ferrari contains 1703 image-caption pairs. [author] Captions contain the names of...
face detection, pedestrian, text, object poseThis data was collected from text ads found on twelve websites that deal with various farm animal related topics. Information from the ad creative and ...
text, classificationThe HTML source of a web page is given. Users looked at each web page and inidated on a 3 point scale (hot medium cold) 50-100 pages per domain. However...
multivariate, text, classificationThis file contains 9 sets of sanitized user data drawn from the command histories of 8 UNIX computer users at Purdue over the course of up to 2 years (U...
text, sequentialThe table below lists the datasets, the YouTube video ID, the amount of samples in each class and the total number of samples per dataset. Dataset --- ...
text, classificationKEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. I...
univariate, classification, clustering, regression, multivariate, text--- By using a tweet crawler, we collect 2000 labelled tweets (1000 positive tweets and 1000 negative ones) on various topics such as: politics ...
text, classificationdataset are derived from the customers reviews in Amazon Commerce Website for authorship identification. Most previous studies conducted the identi...
multivariate, domain-theory, text, classificationThis dataset was constructed by adding elevation information to a 2D road network in North Jutland, Denmark (covering a region of 185 x 135 km^2). Eleva...
regression, text, clustering, sequentialSVHN is a real-world image dataset for developing machine learning and object recognition algorithms with minimal requirement on data preprocessing and ...
urban, real, recognition, text, streetside, world, streetview, classification, detection, numberPart of the problem in using an automated program to discover the unknown target function is to decide how to encode names such that the program can be ...
text, univariate, classificationThis dataset was created for the Paper 'From Group to Individual Labels using Deep Features', Kotzias et. al,. KDD 2015 Please cite the paper if you wan...
text, classificationThe most important feature of this dataset is its simplicity to use and its being well-documented, which can be widely used in various studies of text a...
regression, multivariate, text, classificationCar Reviews ------------ -Full reviews of cars for model-years 2007, 2008, and 2009 -There are about 140-250 cars for each model year -Extracted fields ...
textFrom the original readme file (please consult it for more information): ------------------------- The documents in the Reuters-21578 collection appeared...
text, classificationStyle, Price, Rating, Size, Season, NeckLine, SleeveLength, waiseline, Material, FabricType, Decoration, Pattern, Type, Recommendation are Attributes in...
text, classification, clusteringDocuments are first obtained via a Web search using AMIEI: an integrated platform for delivering enterprise intelligence, developed by AMI Software ([We...
multivariate, text, clustering, sequentialI collected 64 e-mails from DBWorld newsletter and I used them to train different algorithms in order to classify between 'announces of conferences' and...
text, classificationThis is a data set containing 1080 documents of free text business descriptions of Brazilian companies categorized into a subset of 9 categories catalog...
multivariate, text, classificationThe dataset consists of a total of 3600 documents including 600 news/texts from six categories economy, culture-arts, health, politics, sports and tech...
text, classification, clusteringThis dataset contains sentences extracted from user reviews on a given topic. Example topics are performance of Toyota Camry and sound quality of ipod n...
textYouTube Comedy Slam ([Web Link]) is a video discovery experiment running on YouTube's version of labs (called TestTube) for a few months in 2011 and 201...
text, classificationPlease see the README for the details on the data organization, and so on.
multivariate, text, classification, clustering