The two datasets are related to red and white variants of the Portuguese "Vinho Verde" wine. For more details, consult: [Web Link] or the reference [Cortez et al., 2009]. Due to privacy and logistic issues, only physicochemical (inputs) and sensory (the output) variables are available (e.g. there is no data about grape types, wine brand, wine selling price, etc.). These datasets can be viewed as classification or regression tasks. The classes are ordered and not balanced (e.g. there are munch more normal wines than excellent or poor ones). Outlier detection algorithms could be used to detect the few excellent or poor wines. Also, we are not sure if all input variables are relevant. So it could be interesting to test feature selection methods.
KEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. I...
univariate, classification, clustering, regression, multivariate, textThe dataset is composed by two tables. The first table go_track_tracks presents general attributes and each instance has one trajectory that is represen...
regression, multivariate, classificationThis data set contains the acquired time series from 16 chemical sensors exposed to gas mixtures at varying concentration levels. In particular, we gene...
regression, time-series, multivariate, classificationThe experiments have been carried out with a group of 115 students of first-year, undergraduate Engineering major of the University of Genoa. We carri...
classification, clustering, sequential, regression, time-series, multivariateMany real world applications need to know the localization of a user in the world to provide their services. Therefore, automatic user localization has ...
regression, multivariate, classificationA description of the underlying Cargo 2000 standard and the processes reflected in the data set can be found at [Web Link].
regression, multivariate, classification, sequentialKEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. I...
univariate, classification, clustering, regression, multivariate, textCollect the real time readings for residential,commercial,industrial,agriculure,to find the accuracy consumption in Tamil Nadu Around Thanajvur
regression, multivariate, classification, clusteringEach record represents follow-up data for one breast cancer case. These are consecutive patients seen by Dr. Wolberg since 1984, and include only those...
regression, multivariate, classificationThe PD database consists of training and test files. The training data belongs to 20 PWP (6 female, 14 male) and 20 healthy individuals (10 female, 10 m...
regression, multivariate, classificationIndoor localisation is a key topic for the Ambient Intelligence (AmI) research community. In this scenarios, recent advancements in wearable technolog...
classification, clustering, sequential, regression, time-series, multivariateData is collected from imkb.gov.tr and finance.yahoo.com. Data is organized with regard to working days in Istanbul Stock Exchange.
univariate, classification, regression, time-series, multivariateThe dataset was built from a personal collection of 1059 tracks covering 33 countries/area. The music used is traditional, ethnic or `world' only, as c...
regression, multivariate, classificationThe measured data was collected using a chemical sensing system based on an array of 16 metal-oxide gas sensors and an external mechanical ventilator to...
regression, time-series, multivariate, classificationThe main goal of this data set is providing clean and valid signals for designing cuff-less blood pressure estimation algorithms. The raw electrocardiog...
regression, multivariate, classificationIndoor localization is a key topic for mobile computing. However, it is still very difficult for the mobile sensing community to compare state-of-art In...
classification, clustering, sequential, regression, time-series, multivariateA chemical detection platform composed of 8 chemo-resistive gas sensors was exposed to turbulent gas mixtures generated naturally in a wind tunnel. The ...
regression, time-series, multivariate, classificationThe most important feature of this dataset is its simplicity to use and its being well-documented, which can be widely used in various studies of text a...
regression, multivariate, text, classificationThis data set contains 13,910 measurements from 16 chemical sensors exposed to 6 gases at different concentration levels. This dataset is an extension o...
causa, classification, clustering, regression, time-series, multivariateThis dataset includes the recordings of five replicates of an 8-sensor array. Each unit holds 8 MOX sensors and integrates custom-designed electronics f...
classification, regression, time-series, multivariate, domain-theoryProvide all relevant information about your data set.
regression, multivariate, classificationOpen University Learning Analytics Dataset (OULAD) contains data about courses, students and their interactions with Virtual Learning Environment (VLE) ...
classification, clustering, sequential, regression, time-series, multivariateThe PD and control handwriting database consists of 62 PWP (People with parkinson) and 15 healthy individuals who appealed at the Department of Neurolog...
regression, multivariate, classification, clustering* The articles were published by Mashable (www.mashable.com) and their content as the rights to reproduce it belongs to them. Hence, this dataset does n...
regression, multivariate, classificationThe DrivFace database contains images sequences of subjects while driving in real scenarios. It is composed of 606 samples of 640480 pixels each, acquir...
regression, multivariate, classification, clusteringWe perform energy analysis using 12 different building shapes simulated in Ecotect. The buildings differ with respect to the glazing area, the glazing a...
regression, multivariate, classificationThis data approach student achievement in secondary education of two Portuguese schools. The data attributes include student grades, demographic, social...
regression, multivariate, classificationAIMS AND PURPOSES This corpus is intended to do cleaning (or binarization) and enhancement of noisy grayscale printed text images using supervised lear...
regression, multivariate, classificationNotes: -- 3 classes of waves -- 21 attributes, all of which include noise -- See the book for details (49-55, 169) -- waveform.data....
multivariate, generator, classificationThe "spam" concept is diverse: advertisements for products/web sites, make money fast schemes, chain letters, pornography... Our collection of spam e-m...
multivariate, classificationThis dataset contains 503 sponges belonging to the Demospongiae class collected from the Mediterranean (451 sponges) and Atlantic oceans (52 sponges). E...
multivariate, classificationThe REALDISP (REAListic sensor DISPlacement) dataset has been originally collected to investigate the effects of sensor displacement in the activity rec...
time-series, multivariate, classificationThe donation includes 5 datasets, each of them defining a different learning problem: * LP1: failures in approach to grasp position * LP2: fail...
time-series, multivariate, classificationOur dataset is used by us to explore spammers in microblog and you can access our demo system at [Web Link]Please add :8080 after the domain name as por...
univariate, classification, sequential, causal-discovery, multivariate, textTRAINING File : I have created training file with 100+ non malacious examples and 250+ malacious samples. NON-MALACIOUS dataset is represented by +1 whi...
multivariate, classificationThere are four data sets representing different conditions of an experiment. All have the same attributes. a. adult-stretch.data Inflated is true if a...
multivariate, classificationThis data set contains 416 liver patient records and 167 non liver patient records.The data set was collected from north east of Andhra Pradesh, India. ...
multivariate, classificationThe dataset describes diagnosing of cardiac Single Proton Emission Computed Tomography (SPECT) images. Each of the patients is classified into two categ...
multivariate, classificationSamples arrive periodically as Dr. Wolberg reports his clinical cases. The database therefore reflects this chronological grouping of the data. This gro...
multivariate, classificationThis is a tiny database. Michie reports that Burke's group used RULEMASTER to generate comprehendable rules for determining the conditions under which ...
multivariate, classificationThis database contains 76 attributes, but all published experiments refer to using a subset of 14 of them. In particular, the Cleveland database is the...
multivariate, classificationThe datas time period is between Jan 1st, 2010 to Dec 31st, 2014. Missing data are denoted as NA.
regression, time-series, multivariateThis DataSet is about the results of Statlog project. The project performed a comparative study between Statistical, Neural and Symbolic learning algori...
multivariate, classificationThe references below describe a predecessor to this dataset and its development. They also give results (not cross-validated) for classification by a ru...
multivariate, classificationThis data was used by Hong and Young to illustrate the power of the optimal discriminant plane even in ill-posed settings. Applying the KNN method in th...
multivariate, classificationProvide all relevant information about your data set.
multivariate, classification, clusteringThe Infra-Red Astronomy Satellite (IRAS) was the first attempt to map the full sky at infra-red wavelengths. This could not be done from ground observa...
multivariate, classificationThe data is related to posts' published during the year of 2014 on the Facebook's page of a renowned cosmetics brand. This dataset contains 500 of the 7...
regression, multivariateThe data are MC generated (see below) to simulate registration of high energy gamma particles in a ground-based atmospheric Cherenkov gamma telescope us...
multivariate, classificationExtraction was done by Barry Becker from the 1994 Census database. A set of reasonably clean records was extracted using the following conditions: ((AA...
multivariate, classificationThis data set was generated to model psychological experimental results. Each example is classified as having the balance scale tip to the right, tip t...
multivariate, classificationThis data set includes descriptions of hypothetical samples corresponding to 23 species of gilled mushrooms in the Agaricus and Lepiota Family (pp. 500-...
multivariate, classificationThis data set contains weighted census data extracted from the 1994 and 1995 Current Population Surveys conducted by the U.S. Census Bureau. The data co...
multivariate, classificationThe data was collected for examining our newly developed classifier for multidimensional curves (multidimensional time series). Nine male speakers utter...
time-series, multivariate, classificationData were extracted from images that were taken from genuine and forged banknote-like specimens. For digitization, an industrial camera usually used fo...
multivariate, classificationFor a list of attributes, please refer to those two .names files. They use the following naming convention: All the attribute start with T means the t...
time-series, multivariate, classification, sequentialThis dataset comprises information regarding the ADLs performed by two users on a daily basis in their own homes. This dataset is composed by two insta...
classification, clustering, sequential, time-series, multivariateThis dataset contains records of simulation crashes encountered during climate model uncertainty quantification (UQ) ensembles. Ensemble members were ...
multivariate, classificationThe NASA data set comprises different size NACA 0012 airfoils at various wind tunnel speeds and angles of attack. The span of the airfoil and the observ...
regression, multivariateThe user first creates a classification model and then generates classified examples from it. To create a model, the following are specified: the number...
multivariate, classification* The dataset was acquired and annotated by professional physicians at 'Hospital Universitario de Caracas'. * The subjective judgments (target variables...
multivariate, classificationThis simple domain contains 7 Boolean attributes and 10 concepts, the set of decimal digits. Recall that LED displays contain 7 light-emitting diodes -...
multivariate, generator, classificationNews are grouped into clusters that represent pages discussing the same news story. The dataset includes also references to web pages that, at the acce...
multivariate, classification, clusteringNotes: -- 3 classes of waves -- 40 attributes, all of which include noise -- The latter 19 attributes are all noise attributes with me...
multivariate, generator, classificationThis corpus has been collected from free or free for research sources at the Internet: -> A collection of 425 SMS spam messages was manually extracted ...
text, classification, clustering, multivariate, domain-theoryA small subset of the original soybean database. See the reference for Fisher and Schlimmer in soybean-large.names for more information. Steven Souder...
multivariate, classificationData was captured using a setup that consisted of: - Two Fifth Dimension Technologies (5DT) gloves, one right and one left - Two Ascension Flock-of-B...
time-series, multivariate, classificationThe original data were formatted by Thorsten Joachims in the bag-of-words representation. There were 9947 features (of which 2562 are always zeros for a...
multivariate, classificationThis data set contains time series of greenhouse gas (GHG) concentrations at 2921 grid cells in California created using simulations of the Weather Rese...
regression, time-series, multivariateThis database contains 279 attributes, 206 of which are linear valued and the rest are nominal. Concerning the study of H. Altay Guvenir: "The aim is ...
multivariate, classificationThe dataset contains cases from a study that was conducted between 1958 and 1970 at the University of Chicago's Billings Hospital on the survival of pat...
multivariate, classificationMany variables are included so that algorithms that select or learn weights for attributes could be tested. However, clearly unrelated attributes were...
regression, multivariateBiophysical models of mutant p53 proteins yield features which can be used to predict p53 transcriptional activity. All class labels are determined via...
multivariate, classificationThe dataset is the subset of RCV1. These corpus has already been used in author identification experiments. In the top 50 authors (with respect to total...
text, classification, clustering, multivariate, domain-theoryThis data set contains some training and testing data from a remote sensing study by Johnson et al. (2013) that involved detecting diseased trees in Qui...
multivariate, classificationThis data set was generated as follows. 150 subjects spoke the name of each letter of the alphabet twice. Hence, we have 52 training examples from each ...
multivariate, classificationThe digits have been size-normalized and centered in a fixed-size image of dimension 28x28. The original data were modified for the purpose of the featu...
multivariate, classificationThe time period is between Jan 1st, 2010 to Dec 31st, 2015. Missing data are denoted as NA.
regression, time-series, multivariateA Vicon motion capture camera system was used to record 12 users performing 5 hand postures with markers attached to a left-handed glove. A rigid patt...
multivariate, classification, clusteringThis file concerns credit card applications. All attribute names and values have been changed to meaningless symbols to protect confidentiality of the ...
multivariate, classificationWe use the following representation to collect the dataset age - age bp - blood pressure sg - specific gravity al ...
multivariate, classificationThe data set is at 10 min for about 4.5 months. The house temperature and humidity conditions were monitored with a ZigBee wireless sensor network. Each...
regression, time-series, multivariateNotes: -- The database contains 3 potential classes, one for the number of times a certain type of solar flare occured in a 24 hour period. -- Ea...
regression, multivariateMammography is the most effective method for breast cancer screening available today. However, the low positive predictive value of breast biopsy result...
multivariate, classificationThe dataset represents 10 years (1999-2008) of clinical care at 130 US hospitals and integrated delivery networks. It includes over 50 features represen...
multivariate, classification, clusteringThe MHEALTH (Mobile HEALTH) dataset comprises body motion and vital signs recordings for ten volunteers of diverse profile while performing several phys...
time-series, multivariate, classificationPast Usage: (a) Rgnvaldsson, You and Garwicz (2015) 'State of the art prediction of HIV-1 protease cleavage sites', Bioinformatics, vol 31 (8),...
multivariate, classificationSamples (instances) are stored row-wise. Variables (attributes) of each sample are RNA-Seq gene expression levels measured by illumina HiSeq platform.
multivariate, classification, clusteringWe used preprocessing programs made available by NIST to extract normalized bitmaps of handwritten digits from a preprinted form. From a total of 43 peo...
multivariate, classificationThe submitted file is set up as follows. In the first line is the the number of signal events followed by the number of background events. The signal ev...
multivariate, classificationThe examples are complete and noise free. The examples highly simplified the problem. The attributes do not fully describe all the factors affecting the...
multivariate, classification# From Garavan Institute # Documentation: as given by Ross Quinlan # 6 databases from the Garavan Institute in Sydney, Australia # Approximately the fol...
multivariate, domain-theory, classificationFormat: Each observation concerns one university. In some cases, more information is provided about the attribute (e.g., units or domain). Some duplicat...
multivariate, classificationThis dataset was used in several classifications tasks related to the challenge of anuran species recognition through their calls. It is a multilabel da...
multivariate, classification, clusteringBackground information: The data set concerns the earliest history of mankind. Prehistoric men created the desired shape of a stone tool by striking on ...
multivariate, classification, clustering, causal-discoveryThe experiments have been carried out by means of a numerical simulator of a naval vessel (Frigate) characterized by a Gas Turbine (GT) propulsion plant...
regression, multivariateThe source datasets needed to be combined via programming. Many variables are included so that algorithms that select or learn weights for attributes co...
regression, multivariatePredicting forest cover type from cartographic variables only (no remotely sensed data). The actual forest cover type for a given observation (30 x 30 ...
multivariate, classificationThis data set contains training and testing data from a remote sensing study which mapped different forest types based on their spectral characteristics...
multivariate, classificationThese data are the results of a chemical analysis of wines grown in the same region in Italy but derived from three different cultivars. The analysis de...
multivariate, classificationThis is one of three domains provided by the Oncology Institute that has repeatedly appeared in the machine learning literature. (See also lymphography ...
multivariate, classificationTo demonstrate the RFMTC marketing model (a modified version of RFM), this study adopted the donor database of Blood Transfusion Service Center in Hsin-...
multivariate, classificationThis database has been artificially generated by using a first order theory which describes the structure of ten capital letters of the English alphabet...
multivariate, classificationThis archive contains 13910 measurements from 16 chemical sensors utilized in simulations for drift compensation in a discrimination task of 6 gases at ...
multivariate, classificationThis dataset is a subset of the 1987 National Indonesia Contraceptive Prevalence Survey. The samples are married women who were either not pregnant or d...
multivariate, classificationThe data was collected retrospectively at Wroclaw Thoracic Surgery Centre for patients who underwent major lung resections for primary lung cancer in th...
multivariate, classificationBrief Description of the Dataset: --------------------------------- Each of the 19 activities is performed by eight subjects (4 female, 4 male, between ...
time-series, multivariate, classification, clusteringThe PAMAP2 Physical Activity Monitoring dataset contains data of 18 different physical activities (such as walking, cycling, playing soccer, etc.), perf...
time-series, multivariate, classificationThis dataset describes a set of 92 molecules of which 47 are judged by human experts to be musks and the remaining 45 molecules are judged to be non-mus...
multivariate, classificationThis dataset is composed of a range of biomedical voice measurements from 31 people, 23 with Parkinson's disease (PD). Each column in the table is a par...
multivariate, classification2126 fetal cardiotocograms (CTGs) were automatically processed and the respective diagnostic features measured. The CTGs were also classified by three e...
multivariate, classificationNotes: - Additional "background" knowledge is supplied that provides a partial ordering on some of the attribute values. - We are providing this datas...
multivariate, classificationAutomatic identification of commercial blocks in news videos finds a lot of applications in the domain of television broadcast analysis and monitoring. ...
multivariate, classification, clusteringNumber of instances 1030 Number of Attributes 9 Attribute breakdown 8 quantitative input variables, and 1 quantitative output variable Missing Attribut...
regression, multivariateThere are two versions to the database: - V1 contains the original examples and - V2 contains descriptions after discretizing numeric proper...
multivariate, classificationThe file animals.c is a data generator of structured instances representing quadruped animals as used by Gennari, Langley, and Fisher (1989) to evaluate...
multivariate, generator, classificationThis is one of three domains provided by the Oncology Institutenthat has repeatedly appeared in the machine learning literature. (See also breast-cance...
multivariate, classificationIn [Cortez and Morais, 2007], the output 'area' was first transformed with a ln(x+1) function. Then, several Data Mining methods were applied. After ...
regression, multivariatePredicting the age of abalone from physical measurements. The age of abalone is determined by cutting the shell through the cone, staining it, and coun...
multivariate, classificationPlease find the original data at '[Web Link]'
time-series, multivariate, classification, clusteringProvide all relevant information about your data set.
regression, multivariateFor further details on this dataset and/or its attributes, please read the 'ReadMe.pdf' file included and/or consult the Master's Thesis 'Development of...
multivariate, classificationThis database does NOT use a standard set of attributes per instance. Contact Ray Bareiss (rbareiss '@' uunet.uucp ?) for more information. Domain exp...
multivariate, classificationUncompressing rcv1rcv2aminigoutte.tar.bz2 will create a directory that contains 5 subdirectories EN, FR, GR, IT and SP, corresponding to the 5 language...
multivariate, classificationTwo datasets are provided. the original dataset, in the form provided by Prof. Hofmann, contains categorical/symbolic attributes and is in the file "ge...
multivariate, classificationThe measurements were created to ease the development, comparison and evaluation of fingerprinting based hybrid indoor positioning methods. The measurem...
time-series, multivariate, classification, sequentialDatabase contains records for 1885 respondents. For each respondent 12 attributes are known: Personality measurements which include NEO-FFI-R (neurotici...
multivariate, classificationThis dataset contains features extracted from the Messidor image set to predict whether an image contains signs of diabetic retinopathy or not. All feat...
multivariate, classificationA Vicon motion capture camera system was used to record 12 users performing 5 hand postures with markers attached to a left-handed glove. A rigid patt...
multivariate, classification, clusteringThis MALDI-TOF dataset consists in:A) A reference panel of 20 Gram positive and negative bacterial species covering 9 genera among which several species...
multivariate, classificationInformation about customers consists of 86 variables and includes product usage data and socio-demographic data derived from zip area codes. The data wa...
regression, multivariate, descriptionThree data sets are submitted, for training and testing. Ground-truth occupancy was obtained from time stamped pictures that were taken every minute. Fo...
time-series, multivariate, classificationThe data consist of 16 binary inputs and one 'four-bit' one-hot classification output. The 16-bit inputs are binary-valued attack-point vectors. 1 indic...
multivariate, classificationThe dataset is about bankruptcy prediction of Polish companies. The data was collected from Emerging Markets Information Service (EMIS, [Web Link]), whi...
multivariate, classificationThis dataset represents a real-life benchmark in the area of Activity Recognition applications, as described in [1]. The classification tasks consist ...
time-series, multivariate, classification, sequentialThere are two databases: (both use the same set of 5 attributes): 1. Primary o-ring erosion and/or blowby 2. Primary o-ring erosion only The two databa...
regression, multivariateCar Evaluation Database was derived from a simple hierarchical decision model originally developed for the demonstration of DEX, M. Bohanec, V. Rajkovic...
multivariate, classificationThe experiments have been carried out with a group of 30 volunteers within an age bracket of 19-48 years. Each person performed six activities (WALKING,...
time-series, multivariate, classification, clusteringThe Daphnet Freezing of Gait Dataset is a dataset devised to benchmark automatic methods to recognize gait freeze from wearable acceleration sensors pl...
time-series, multivariate, classificationThis database encodes the complete set of possible board configurations at the end of tic-tac-toe games, where "x" is assumed to have played first. The...
multivariate, classificationThere are 19 classes, only the first 15 of which have been used in prior work. The folklore seems to be that the last four classes are unjustified by th...
multivariate, classificationThis dataset was derived from geospatial data from two sources: 1) Landsat time-series satellite imagery from the years 2014-2015, and 2) crowdsourced g...
multivariate, classificationThe data consist of evaluations of teaching performance over three regular semesters and two summer semesters of 151 teaching assistant (TA) assignments...
multivariate, classificationF. Bergadano supplied this database. Each instance contains many components, each of which has 8 attributes. Different instances in this database have...
multivariate, classificationA dataset of online handwritten assamese characters by collecting samples from 45 writers is created. Each writer contributed 52 basic characters, 10 nu...
multivariate, classification, sequentialARCENE was obtained by merging three mass-spectrometry datasets to obtain enough training and test data for a benchmark. The original features indicate ...
multivariate, classificationThis dataset is composed of a range of biomedical voice measurements from 42 people with early-stage Parkinson's disease recruited to a six-month trial ...
regression, multivariateThe phishing problem is considered a vital issue in .COM industry especially e-banking and e-commerce taking the number of online transactions involving...
multivariate, classificationWe have downloaded 15 months worth of daily data from the California Department of Transportation PEMS website, [Web Link], The data describes the occup...
time-series, multivariate, classificationThe companion file is a Common Lisp demonstration file that generates knight-pin Chess end-game samples. Start up Lisp and load the file. It generates ...
multivariate, generator, classificationThis dataset is a slightly modified version of the dataset provided in the StatLib library. In line with the use by Ross Quinlan (1993) in predicting t...
regression, multivariateThis research aimed at the case of customers default payments in Taiwan and compares the predictive accuracy of probability of default among six data mi...
multivariate, classificationSee the file bridge-holden-paulson-details.txt in the submitted tarball.
multivariate, classificationThe parameters which we used for collecting the dataset is referred from the paper 'The discovery of experts decision rules from qualitative bankruptcy ...
multivariate, classificationFeatures are computed from a digitized image of a fine needle aspirate (FNA) of a breast mass. They describe characteristics of the cell nuclei present...
multivariate, classificationThe purpose is to classify a given silhouette as one of four types of vehicle, using a set of features extracted from the silhouette. The vehicle may b...
multivariate, classificationThe data set includes 103 data points. There are 7 input variables, and 3 output variables in the data set. The initial data set included 78 data. After...
regression, multivariateOngoing research on university faculty perceptions and practices of using Wikipedia as a teaching resource. Based on a Technology Acceptance Model, the ...
regression, multivariate, clustering, causal-discovery-- We aggregated screen movements into screen-fixations using a Salvucci & Goldberg (2000) dispersion-threshold algorithm, and defined Perception Act...
regression, multivariateContains training and testing data for classifying a high resolution aerial image into 9 types of urban land cover. Multi-scale spectral, size, shape, a...
multivariate, classificationThe presented dataset is composed of two tsv files named 'youtube_videos.tsv' and 'transcoding_mesurment.tsv'. The first contains 10 columns of fundame...
regression, multivariateThe file "sonar.mines" contains 111 patterns obtained by bouncing sonar signals off a metal cylinder at various angles and under various conditions. Th...
multivariate, classificationThe data set consists of the expression levels of 77 proteins/protein modifications that produced detectable signals in the nuclear fraction of cortex. ...
multivariate, classification, clusteringVelloso, E.; Bulling, A.; Gellersen, H.; Ugulino, W.; Fuks, H. Qualitative Activity Recognition of Weight Lifting Exercises. Proceedings of 4th Internat...
multivariate, classificationdataset are derived from the customers reviews in Amazon Commerce Website for authorship identification. Most previous studies conducted the identi...
multivariate, domain-theory, text, classificationThis dataset describes a set of 102 molecules of which 39 are judged by human experts to be musks and the remaining 63 molecules are judged to be non-mu...
multivariate, classificationThe dataset describes diagnosing of cardiac Single Proton Emission Computed Tomography (SPECT) images. Each of the patients is classified into two categ...
multivariate, classificationThe dataset (movement_libras) contains 15 classes of 24 instances each, where each class references to a hand movement type in LIBRAS. In the video pre...
multivariate, classification, clustering, sequentialThe estimated relative performance values were estimated by the authors using a linear regression method. See their article (pp 308-313) for more detai...
regression, multivariateWe have created the UJIpenchars2 character database by collecting samples from 60 writers at two different sites in two phases: 1st phase, 11 writers, c...
multivariate, classification, sequentialAll data is from one continuous EEG measurement with the Emotiv EEG Neuroheadset. The duration of the measurement was 117 seconds. The eye state was det...
time-series, multivariate, classification, sequentialPrediction of residuary resistance of sailing yachts at the initial design stage is of a great value for evaluating the ships performance and for estima...
regression, multivariateThis database contains all legal 8-ply positions in the game of connect-4 in which neither player has won yet, and in which the next move is not forced....
multivariate, spatial, classificationPlease see the README for the details on the data organization, and so on.
multivariate, text, classification, clusteringThis database contains 34 attributes, 33 of which are linear valued and one of them is nominal. The differential diagnosis of erythemato-squamous dise...
multivariate, classificationAll the patients suffered heart attacks at some point in the past. Some are still alive and some are not. The survival and still-alive variables, when ...
multivariate, classificationThis file concerns credit card applications. All attribute names and values have been changed to meaningless symbols to protect confidentiality of the ...
multivariate, classificationUncompressing the archive url_svmlight.tar.gz will yield a directory url_svmlight/ containing the following files: * FeatureTypes --- A text file li...
time-series, multivariate, classificationType of dependent variables (7 Types of Steel Plates Faults): 1.Pastry 2.Z_Scratch 3.K_Scatch 4.Stains 5.Dirtiness 6.Bumps 7.Other_Faults
multivariate, classificationThe examined group comprised kernels belonging to three different varieties of wheat: Kama, Rosa and Canadian, 70 elements each, randomly selected for t...
multivariate, classification, clusteringDota 2 is a popular computer game with two teams of 5 players. At the start of the game each player chooses a unique hero with different strengths and w...
multivariate, classificationFeatures are extracted from electric current drive signals. The drive has intact and defective components. This results in 11 different classes with dif...
multivariate, classification-- The users' knowledge class were classified by the authors using intuitive knowledge classifier (a hybrid ML technique of k-NN and meta-heurist...
multivariate, classification, clusteringThis data set includes votes for each of the U.S. House of Representatives Congressmen on the 16 key votes identified by the CQA. The CQA lists nine di...
multivariate, classificationPredicted Attribute: Localization site of protein. ( non-numeric ). The references below describe a predecessor to this dataset and its development. Th...
multivariate, classification-- The users' knowledge class were classified by the authors using intuitive knowledge classifier (a hybrid ML technique of k-NN and meta-heuristi...
multivariate, classificationBiomedical data set built by Dr. Henrique da Mota during a medical residence period in the Group of Applied Research in Orthopaedics (GARO) of the Centr...
multivariate, classification2 data files: -- horse-colic.data: 300 training instances -- horse-colic.test: 68 test instances Possible class attributes: 24 (whether le...
multivariate, classificationThe MONK's problem were the basis of a first international comparison of learning algorithms. The result of this comparison is summarized in "The MONK's...
multivariate, classificationThe dataset is composed by features extracted from 7 videos with people gesticulating, aiming at studying Gesture Phase Segmentation. Each video is repr...
classification, clustering, sequential, time-series, multivariateThe dataset could contain missing values. The data was sampled every minute, computing and uploading it smoothed with 15 minute means. The header of the...
text, sequential, regression, time-series, multivariateThe QSAR biodegradation dataset was built in the Milano Chemometrics and QSAR Research Group (Universit degli Studi Milano Bicocca, Milano, Italy). The...
multivariate, classificationApproximately 80% of the data belongs to class 1. Therefore the default accuracy is about 80%. The aim here is to obtain an accuracy of 99 - 99.9%. The...
multivariate, classification* Audio track (encoded as mp3) of each of the 106,574 tracks. It is on average 10 millions samples per track.* Nine audio features (consisting of 518 at...
time-series, multivariate, classification, clusteringThis dataset represents a set of possible advertisements on Internet pages. The features encode the geometry of the image (if available) as well as phr...
multivariate, classificationWe create a character database by collecting samples from 11 writers. Each writer contributed with letters (lower and uppercase), digits, and other cha...
multivariate, classification, sequentialThe HTML source of a web page is given. Users looked at each web page and inidated on a 3 point scale (hot medium cold) 50-100 pages per domain. However...
multivariate, text, classificationVina conducted a comparison test of her rule-based system, BEAGLE, the nearest-neighbor algorithm, and discriminant analysis. BEAGLE is a product avail...
multivariate, classificationThe OPPORTUNITY Dataset for Human Activity Recognition from Wearable, Object, and Ambient Sensors is a dataset devised to benchmark human activity recog...
time-series, multivariate, classificationMADELON is an artificial dataset containing data points grouped in 32 clusters placed on the vertices of a five dimensional hypercube and randomly label...
multivariate, classificationThe dataset contains 9358 instances of hourly averaged responses from an array of 5 metal oxide chemical sensors embedded in an Air Quality Chemical Mul...
regression, time-series, multivariateExamples represent positive and negative instances of people who were and were not granted credit. The theory was generated by talking to the individu...
multivariate, domain-theory, classificationThis is perhaps the best known database to be found in the pattern recognition literature. Fisher's paper is a classic in the field and is referenced f...
multivariate, classificationThe objective is to identify each of a large number of black-and-white rectangular pixel displays as one of the 26 capital letters in the English alphab...
multivariate, classificationThe dataset format is described below. Note: the format of this database was modified on 2/26/90 to conform with the format of all the other databases ...
multivariate, classificationImpedance measurements were made at the frequencies: 15.625, 31.25, 62.5, 125, 250, 500, 1000 KHz Impedance measurements of freshly excised breast tissu...
multivariate, classificationThere are three disadvantages of weighted scoring stock selection models. First, they cannot identify the relations between weights of stock-picking con...
regression, multivariateThis dataset represents a real-life benchmark in the area of Ambient Assisted Living applications, as described in [1]. The binary classification task c...
time-series, multivariate, classification, sequential21 bioassay datasets generated from Pubchem. Both Primary and confirmatory bioassays (12 bioassays, 21 mixes)The data is provided in the same train/test...
multivariate, classificationThis database contains 5 numeric-valued attributes. Only a subset of 3 are used during testing (the latter 3). Furthermore, only 2 of the 3 concepts a...
multivariate, classificationNorthix is designed to be a schema matching benchmark problem for data integration of two entity relationship databases. Northix is the resulting schema...
multivariate, text, univariate, classificationThis dataset contains the features extracted from a database of colonoscopic videos showing gastrointestinal lesions. It also contains the ground truth ...
multivariate, classificationThe original paper demonstrated that it is possible to correctly replicate the experts' binary assessment with approximately 90% accuracy using both 10-...
multivariate, classificationAll the 504 reviews were collected between January and August of 2015.
regression, classification1593 handwritten digits from around 80 persons were scanned, stretched in a rectangular box 16x16 in a gray scale of 256 values.Then each pixel of each ...
multivariate, classificationThe main idea of this data set is to prepare the algorithm of the expert system, which will perform the presumptive diagnosis of two diseases of urinar...
multivariate, classificationThis archive contains 2075259 measurements gathered between December 2006 and November 2010 (47 months). Notes: 1.(global_active_power*1000/60 - sub_me...
regression, time-series, multivariate, clusteringThis is one of three domains provided by the Oncology Institute that has repeatedly appeared in the machine learning literature. (See also breast-cancer...
multivariate, classificationA simple database containing 17 Boolean-valued attributes. The "type" attribute appears to be the class attribute. Here is a breakdown of which animal...
multivariate, classificationThis database is a standardized version of the original audiology database (see audiology.* in this directory). The non-standard set of attributes have...
multivariate, classificationThe instances were drawn randomly from a database of 7 outdoor images. The images were handsegmented to create a classification for every pixel. Eac...
multivariate, classificationEach record is an example of a hand consisting of five playing cards drawn from a standard deck of 52. Each card is described using two attributes (suit...
multivariate, classificationRoss Quinlan: This data was given to me by Karl Ulrich at MIT in 1986. I didn't record his description at the time, but here's his subsequent (1992) r...
regression, multivariateThe data is related with direct marketing campaigns of a Portuguese banking institution. The marketing campaigns were based on phone calls. Often, more ...
multivariate, classificationThe source of the data is the raw measurements from a Nintendo PowerGlove. It was interfaced through a PowerGlove Serial Interface to a Silicon Graphics...
time-series, multivariate, classificationThis is a data set containing 1080 documents of free text business descriptions of Brazilian companies categorized into a subset of 9 categories catalog...
multivariate, text, classificationNursery Database was derived from a hierarchical decision model originally developed to rank applications for nursery schools. It was used during severa...
multivariate, classificationThe Dataset for ADL Recognition with Wrist-worn Accelerometer is a public collection of labelled accelerometer data recordings to be used for the creati...
time-series, multivariate, classification, clusteringMining activity was and is always connected with the occurrence of dangers which are commonly called mining hazards. A special case of such threat is a...
multivariate, classificationYou should respect the following train / test split: train: first 463,715 examples test: last 51,630 examples It avoids the 'producer effect' by making ...
regression, multivariateThe dataset is used in the evaluation of EER, FRR and FAR metrics using a new anomaly detector model (Med-Min-Diff). The typed text in the experiment is...
multivariate, classificationThe classification task of this database is to determine where patients in a postoperative recovery area should be sent to next. Because hypothermia is...
multivariate, classificationThe dataset contains 9568 data points collected from a Combined Cycle Power Plant over 6 years (2006-2011), when the power plant was set to work with fu...
regression, multivariateThis data file contains details of various nations and their flags. In this file the fields are separated by spaces (not commas). With this data you ca...
multivariate, classificationA complex modern semi-conductor manufacturing process is normally under consistent surveillance via the monitoring of signals/variables collected from s...
multivariate, classification, causal-discoveryThe automated analysis of facial expressions has been widely used in different research areas, such as biometrics or emotional analysis. Special impo...
multivariate, classification, clustering, sequentialThe 5473 examples comes from 54 distinct documents. Each observation concerns one block. All attributes are numeric. Data are in a format readable by C4.5.
multivariate, classificationIn this paper, we look for to recognize the causes of users tend to cyber space in Kohkiloye and Boyer Ahmad Province in Iran. Collecting information to...
multivariate, classificationSeveral constraints were placed on the selection of these instances from a larger database. In particular, all patients here are females at least 21 ye...
multivariate, classificationThe experiments were carried out with a group of 30 volunteers within an age bracket of 19-48 years. They performed a protocol of activities composed of...
time-series, multivariate, classificationThis is a transnational data set which contains all the transactions occurring between 01/12/2010 and 09/12/2011 for a UK-based and registered non-store...
classification, clustering, sequential, time-series, multivariateWe create a digit database by collecting 250 samples from 44 writers. The samples written by 30 writers are used for training, cross-validation and writ...
multivariate, classificationThis dataset has recordings of a gas sensor array composed of 8 MOX gas sensors, and a temperature and humidity sensor. This sensor array was exposed to...
time-series, multivariate, classificationThe dataset was collected at 'Hospital Universitario de Caracas' in Caracas, Venezuela. The dataset comprises demographic information, habits, and histo...
multivariate, classificationThis radar data was collected by a system in Goose Bay, Labrador. This system consists of a phased array of 16 high-frequency antennas with a total tra...
multivariate, classificationDrugs are typically small organic molecules that achieve their desired activity by binding to a target site on a receptor. The first step in the discove...
multivariate, classificationDataset from 8800(10 digits x 10 repetitions x 88 speakers) time series of 13 Frequency Cepstral Coefficients (MFCCs) had taken from 44 males and 44 fem...
time-series, multivariate, classificationThe Dataset is uploaded in ZIP format. The dataset contains 5 variants of the dataset, for the details about the variants and detailed analysis read and...
regression, multivariate- The leaves were placed on a white background and then photographed. - The pictures were taken in broad daylight to ensure optimum light intensity.
multivariate, classification, clusteringThis data set consists of three types of entities: (a) the specification of an auto in terms of various characteristics, (b) its assigned insurance risk...
regression, multivariateNumber of instances: 18000 times-series measurements recorded from a 72 metal-oxide gas sensor array-based chemical detection platform. Number of attri...
time-series, multivariate, classificationHTRU2 is a data set which describes a sample of pulsar candidates collected during the High Time Resolution Universe Survey (South) [1]. Pulsars are a...
multivariate, classification, clusteringWe present a dataset to address the problem of visual privacy - where users unintentionally leak private information when sharing personal images online...
multilabel, privacy, classification, flickr, scene, regressionThe database consists of the multi-spectral values of pixels in 3x3 neighbourhoods in a satellite image, and the classification associated with the cent...
multivariate, classificationThis dataset consists of features of handwritten numerals (`0'--`9') extracted from a collection of Dutch utility maps. 200 patterns per class (for a to...
multivariate, classificationThe Heterogeneity Dataset for Human Activity Recognition from Smartphone and Smartwatch sensors consists of two datasets devised to investigate sensor h...
time-series, multivariate, classification, clusteringHere's the abstract from the above reference: ABSTRACT: Machine learning tools show significant promise for knowledge acquisition, particularly when hu...
multivariate, classificationThe provided files comprise three different data sets. The first one contains the raw values of the measurements of all 24 ultrasound sensors and the c...
multivariate, classification, sequentialThe records represent individual data including first and family name, sex, date of birth and postal code, which were collected through iterative insert...
multivariate, classificationThe instances were drawn randomly from a database of 7 outdoor images. The images were handsegmented to create a classification for every pixel. ...
multivariate, classificationMachine learning is used in high-energy physics experiments to search for the signatures of exotic particles. These signatures are learned from Monte Ca...
multivariate, classificationExtraction was done by Barry Becker from the 1994 Census database. A set of reasonably clean records was extracted using the following conditions: ((AA...
multivariate, classificationPlease ask Gail Gong for further information on this database.
multivariate, classificationCost Matrix _______ abse pres absence 0 1 presence 5 0 where the rows represent the true values and the columns the predicted.
multivariate, classificationAn Inductive Logic Programming (ILP) or relational learning framework is assumed (Muggleton, 1992). The learning system is provided with examples of che...
multivariate, classificationThis data originates from blog posts. The raw HTML-documents of the blog posts were crawled and processed. The prediction task associated with the dat...
regression, multivariateDataset A (former NLPR Gait Database) was created on Dec. 10, 2001, including 20 persons. Each person has 12 image sequences, 4 sequences for each of th...
motion, foot, human, recognition, gait, action, classification, biometry, pressureProblem Description: Splice junctions are points on a DNA sequence at which `superfluous' DNA is removed during the process of protein creation in hig...
domain-theory, classification, sequentialWordNet is a large lexical database of English. Nouns, verbs, adjectives and adverbs are grouped into sets of cognitive synonyms (synsets), each express...
language, category, classification, imagenet, hierarchyObservations come from 2 data streams (people flow in and out of the building), over 15 weeks, 48 time slices per day (half hour count aggregates). T...
time-series, multivariatePitch classes information has been extracted from MIDI sources downloaded from (JSB Chorales)[[Web Link]]. Meter information has been computed through t...
classification, sequentialThis is a sparse data set, less than 10% of the attributes are used for each sample. The link is to a '*.tgz' file which contains two files: [amzn-anon-...
clustering, causal-discovery, regression, time-series, domain-theoryThe data set gathered when we were working at project for Bahrain university between 2002 and 2003.
domain-theory, univariate, classification, clusteringDocuments are first obtained via a Web search using AMIEI: an integrated platform for delivering enterprise intelligence, developed by AMI Software ([We...
multivariate, text, clustering, sequentialKNAPSACK_01 is a dataset directory which contains some examples of data for 01 Knapsack problems. In the 01 Knapsack problem, we are given a knapsack ...
classification, machine learning--- By using a tweet crawler, we collect 2000 labelled tweets (1000 positive tweets and 1000 negative ones) on various topics such as: politics ...
text, classificationThe CVC Partial Occlusion Virtual Pedestrian datasets (CVC-01 to CVC-06) cover a range of scenarios of occluded pedestrians generated in a virtual and r...
urban, pedestrian, classification, synthetic, occlusion, tracking, detectionClassification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
detection, classificationAt Udacity, we believe in democratizing education. How can we provide opportunity to everyone on the planet? We also believe in teaching really amazing ...
driving, street, urban, time, recognition, autonomous, video, segmentation, robot, classification, detection, car, syntheticThis dataset has been developed to help evaluate a "hybrid" learning algorithm ("KBANN") that uses examples to inductively refine preexisting knowledge....
domain-theory, classification, sequentialInstrumentation: The data were collected at a sampling rate of 500 Hz, using as a programming kernel the National Instruments (NI) Labview. The signals ...
time-series, classificationPictures of objects belonging to 256 categoriesPictures of objects belonging to 256 categories.
natural-image, classificationThe Fish4Knowledge project (groups.inf.ed.ac.uk/f4k/) is pleased to announce the availability of 2 subsets of our tropical coral reef fish video and e...
motion, nature, recognition, fish, video, water, classification, animal, cameraData Type: GrayScale Image The image dataset can be used to benchmark classification algorithm for OCR systems. The highest accuracy obtained in the Te...
classificationThe problem is specified by the accompanying data file, "vowel.data". This consists of a three dimensional array: voweldata [speaker, vowel, input]. Th...
classificationThis dataset comes from the daily measures of sensors in a urban waste water treatment plant. The objective is to classify the operational state of the ...
multivariate, clusteringThe GaTech VideoContext dataset consists of over 100 groundtruth annotated outdoor videos with over 20000 frames for the task of geometric context eval...
urban, nature, outdoor, video, segmentation, supervised, classification, context, unsupervised, geometry, semanticThis dataset was created for the Paper 'From Group to Individual Labels using Deep Features', Kotzias et. al,. KDD 2015 Please cite the paper if you wan...
text, classificationThe skin dataset is collected by randomly sampling B,G,R values from face images of various age groups (young, middle, and old), race groups (white, bla...
univariate, classificationCSV format where each row is a paper and each column is an attribute.
multivariate, clusteringThis ETHZ CVL RueMonge 2014 dataset used for 3D reconstruction and semantic mesh labelling for urban scene understanding. It was first published in [1...
benchmark, paris, reconstruction, pointcloud, outdoor, 3d, source, architecture, semantic, code, urban, mesh, recognition, segmentation, classificationThe Yotta dataset consists of 70 images for semantic labeling given in 11 classes. It also contains multiple videos and camera matrices for 14km or driv...
urban, reconstruction, video, segmentation, 3d, classification, camera, semanticThe Brodatz dataset consists of 112 textures in grayscale images of various texture types. http://www.ee.oulu.fi/research/imag/texture/image_data/Brod...
segmentation, benchmark, classification, synthetic, textureThe dataset consists of a total of 3600 documents including 600 news/texts from six categories economy, culture-arts, health, politics, sports and tech...
text, classification, clusteringThe first 5 variables are all blood tests which are thought to be sensitive to liver disorders that might arise from excessive alcohol consumption. Each...
multivariateThis dataset provides a collection of web images and 3D models for research on landmark recognition (especially for methods based on 3D models). We hope...
codebook, reconstruction, matching, recognition, retrieval, 3d, classification, feature, flickr, landmarkProvide all relevant informatioThe data has been produced using Monte Carlo simulations. The first 8 features are kinematic properties measured by the p...
classificationA large dataset of geotagged face images collected from Flickr. The zip file contains text files containing urls of the images. Face2GPS: Estimating G...
gender, face, geotagged, classification, age, localization, humanThe data can be used to try to predict student learning in SE teamwork based on observation of their team activity **** README FILE from the submitted...
time-series, classification, sequentialFor Further information about the variables see the file in the data folder.
text, classificationA dataset for testing object class detection algorithms. It contains 255 test images and features five diverse shape-based classes (apple logos, bottles...
classificationThis data arises from a large study to examine EEG correlates of genetic predisposition to alcoholism. It contains measurements from 64 electrodes place...
time-series, multivariateIMPORTANT: we have lower performance on 'leave-one-subject-out' tests. The performance baseline index we established is for 10-fold cross-validation tes...
classification, sequentialThe data was collected as part of the 1990 census. There are 68 categorical attributes. This data set was derived from the USCensus1990raw data set. T...
multivariate, clusteringThe Outex dataset is part of a framework for empirical evaluation of texture classification and segmentation algorithms. The framework is being constru...
segmentation, benchmark, classification, synthetic, textureThe table below lists the datasets, the YouTube video ID, the amount of samples in each class and the total number of samples per dataset. Dataset --- ...
text, classificationThe Caltech Game Covers dataset consists of CD/DVD covers of video games. The set was downloaded from freecovers.net during the summer of 2008. The set ...
caltech, retrieval, game, cover, classification, hierarchy, taxonomyThe CMP Facade dataset consists of facade images assembled at the Center for Machine Perception, which includes 600 rectified images of facades from var...
urban, similarity, facade, recognition, segmentation, structure, classification, rectification, semanticSVHN is a real-world image dataset for developing machine learning and object recognition algorithms with minimal requirement on data preprocessing and ...
urban, real, recognition, text, streetside, world, streetview, classification, detection, numberBelgiumTSC dataset is built for traffic sign classification purposes. Is is a subset of BelgiumTS dataset and contains cropped images around annotations...
urban, traffic, road, classification, sign, belgium1. Protocol: Three male and one female subjects (age 25 to 30), who have experienced aggression in scenarios such as physical fighting, took part ...
time-series, classificationI collected 64 e-mails from DBWorld newsletter and I used them to train different algorithms in order to classify between 'announces of conferences' and...
text, classificationEach record represents 100 points on a two-dimensional graph. When plotted in order (from 1 through 100) as the Y co-ordinate, the points will create ei...
classification, sequentialAbstract Scene understanding has (again) become a focus of computer vision research, leveraging advances in detection, context modeling, and tracking. ...
scene, segmentation, pedestrian, 3d, classification, understanding, car, semanticThis is an updated and corrected version of the data set used by Sejnowski and Rosenberg in their influential study of speech generation using a neural ...
multivariateA dataset for Attribute Based Classification. It consists of 30475 images of 50 animals classes with six pre-extracted feature representations for each ...
classificationThis dataset was collected for training and validation of machine learning algorithm for classification regions of documents on text, picture and backgr...
classificationThe Street View Text (SVT) dataset contains 647 words and 3796 letters in 249 images harvested from Google Street View. The dataset is more challengin...
urban, text recognition, text detection, classification, outdoorBike sharing systems are new generation of traditional bike rentals where whole process from membership, rental and return back has become automatic. Th...
regression, univariateThe data were collected over a series of specifically designed trials. Our hope was to cover most of the types of sensory interactions that a Pioneer mi...
time-series, multivariateStyle, Price, Rating, Size, Season, NeckLine, SleeveLength, waiseline, Material, FabricType, Decoration, Pattern, Type, Recommendation are Attributes in...
text, classification, clusteringUSPTO Algorithm Challenge, run by NASA-Harvard Tournament Lab and TopCoder Problem: Patent Labeling
domain-theory, classificationThe original source for this data set is the IPUMS project (RugglesSobek, 1997). The IPUMS project is a large collection of federal census data which ha...
multivariateThese are atlantic-mediterranean marine sponges that belong to O.Hadromerida (Demospongiae.Porifera).
multivariate, clusteringThis material is supplementary to Michael Stark, Bernt Schiele. How Good are Local Features for Classes of Geometric Objects. Eleventh IEEE Internat...
object, binary, tool, classification, shapeThe Stanford Dogs dataset contains images of 120 breeds of dogs from around the world. This dataset has been built using images and annotation from Imag...
fine-grained categorization, dogs, detection, classificationThe dataset has been enriched during the Nomao Challenge: [Web Link] organized along with the ALRA workshop (Active Learning in Real-world Applications)...
univariate, classificationThe Comprehensive Cars (CompCars) dataset contains data from two scenarios, including images from web-nature and surveillance-nature. The web-nature dat...
object, urban, fine-grained, classification, recognition, vehicle, car, attributeThe Text and Vision (TVGraz) dataset is an annotated multi-modal dataset which currently contains 10 visual object categories, 4030 images and associate...
text, evaluation, appearance, classificationAWS hosts a variety of public datasets that anyone can access for free. Previously, large datasets such as satellite imagery or genomic data have requi...
space, human, recognition, image, amazon, satellite, segmentation, learning, deep, classification, biology, resolutionThe Extreme Classification Repository: Multi-label Datasets & Code Kush Bhatia Himanshu Jain Prateek Jain Manik Varma The objective in extreme mu...
multilabel, machine, learning, benchmark, evaluation, classificationData Characteristics: -------------------- This data was created by selecting 20 files each from the 10 largest classes in the Reuters-21578 collection...
text, classificationThe HandNet dataset contains depth images of 10 participants hands non-rigidly deforming infront of a RealSense RGB-D camera. This dataset includes 2...
rgbd, hand, articulation, video, segmentation, classification, pose, fingertip, detectionThis loop sensor data was collected for the Glendale on ramp for the 101 North freeway in Los Angeles. It is close enough to the stadium to see unusual...
time-series, multivariateVideo data sets to train machines to recognise objects in our environment. e-VDS35 has 35 classes and a total of 2050 videos of roughly 10 seconds each.
classificationEach image can be characterized by the pose, expression, eyes, and size. There are 32 images for each person capturing every combination of features. ...
image, classificationThe objective is to determine the set of boolean rules that describe the interactions of the nodes within this plant signaling network. The dataset incl...
multivariate, causal-discoveryThe dataset captures 25 people preparing 2 mixed salads each and contains over 4h of annotated accelerometer and RGB-D video data. Annotated activities ...
video, activity, classification, tracking, recognition, detection, actionThe original image collection was obtained from Corel at [Web Link]. There are 68,040 photo images from various categories. Each set of features is sto...
multivariateData set has no missing values. Values are in kW of each 15 min. To convert values in kWh values must be divided by 4. Each column represent one client....
regression, time-series, clusteringThe Stanford Background Dataset is a new dataset introduced in Gould et al. (ICCV 2009) for evaluating methods for geometric and semantic scene understa...
segmentation, urban, geometry, semantic, classification, nature2. Information database: 2.1. Protocol: 22 male subjects , 11 with different knee abnormalities previously diagnosed by a professional. They undergo th...
time-series, multivariateThis dataset was constructed by adding elevation information to a 2D road network in North Jutland, Denmark (covering a region of 185 x 135 km^2). Eleva...
regression, text, clustering, sequentialIn SIFT10M, the titles of the png files indicate the columns position of the SIFT features. This data set has been used for evaluating the approximate n...
multivariate, causal-discoveryThis database contains 18000 video frames of 640x480 resolution from 60 video sequences, each of which recorded from a different subject (31 female and ...
classificationPeople used for recording of the data were wearing four tags (ankle left, ankle right, belt and chest). Each instance is a localization data for one of...
time-series, univariate, classification, sequentialThe data sets we propose to analyse are constituted of 1024 vectors, each vector includes 10 parameters. You can think of it as a 1024*10 matrix. To pro...
multivariateThis data comes from a water quality study where samples were taken from sites on different European rivers of a period of approximately one year. These...
multivariateFine-Grained Visual Classification of Aircraft (FGVC-Aircraft) is a benchmark dataset for the fine grained visual categorization of aircraft. Data, an...
benchmark, evaluation, fine-grained, classification, aircraft, airplane, recognitionYouTube Comedy Slam ([Web Link]) is a video discovery experiment running on YouTube's version of labs (called TestTube) for a few months in 2011 and 201...
text, classificationThe data is stored in relational form across several files. The central file (MAIN) is a list of movies, each with a unique identifier. These identifier...
multivariate, relationalThis dataset contains 600 examples of control charts synthetically generated by the process in Alcock and Manolopoulos (1999). There are six different c...
time-series, classification, clusteringThe Webcam Interestingness dataset consists of 20 different webcam streams, with 159 images each. It is annotated with interestingness ground truth, acq...
video, interest, retrieval, classification, weather, ranking, webcamWe created this data by sampling and processing the www.kelkoo.com logs. The data records offers which were clicked (or shown) to the users of the www.k...
multivariate, causal-discoveryThe data was collected by the Magellan spacecraft over an approximately four year period from 1990--1994. The objective of the mission was to obtain glo...
image, classificationEEG record contains many regular oscillations, which are believed to reflect synchronized rhythmic activity in a group of neurons. Most activity related...
univariate, classificationIn predicting stock prices you collect data over some period of time - day, week, month, etc. But you cannot take advantage of data from a time period u...
time-series, classification, clusteringThis dataset contains Australian legal cases from the Federal Court of Australia (FCA). The cases were downloaded from AustLII ([Web Link]). We included...
text, classificationThis is the first smartphone based lifelogging dataset that is going to be available for public use. Please consider that the user of this dataset are o...
multivariate, causal-discoveryThis challenge is set up around three tasks: Text Localisation, Text Segmentation and Word Recognition. Participation in any or all tasks is welcome. Ch...
text recognition, text detection, classificationThe RGB-D Person Re-identification dataset is for person re-identification using depth information. The main motivation is that the standard techniques ...
pedestrian, 3d, identification, classification, depth, shapeStroke Width Transform Text dataset is by Boris Epstein and consists of 307 images and XXX text instances. Detecting Text in Natural Scenes with Stro...
text recognition, text detection, classificationThe ICDAR 2003 datasets available for download on this site: Robust Reading , Robust Word Recognition , Robust OCR , Text Locating and Cursive Script . ...
text recognition, text detection, classificationThis dataset is an addition to the dataset at [Web Link] We collected more dataset to improve the accuracy of our HAR algorithms applied in ...
time-series, classificationThe measurements were created to ease the development, comparison and evaluation of fingerprinting based hybrid indoor positioning methods. The measurem...
text, classification, clustering, causal-discovery--- The dataset collects data from a wearable accelerometer mounted on the chest --- Sampling frequency of the accelerometer: 52 Hz --- Acceler...
univariate, classification, clustering, sequential, time-series1. Title of Database: Machine Learning based ZZAlpha Stock Recommendations 2. Sources: (a) Original owners of data: ZZAlpha Ltd., 4729 E. Sunrise #10...
time-series, classification, sequentialThe Daimler Mono Pedestrian Classification Benchmark dataset consists of two parts: a base data set. The base data set contains a total of 4000 pedest...
illumination, object, urban, pedestrian, classification, outdoor, scaleThe Pittsburgh Fast-food Image dataset (PFID) consists of 4545 still images, 606 stereo pairs, 3033600 videos for structure from motion, and 27 privacy-...
video, laboratory, classification, reconstruction, real, food, recognitionDiabetes patient records were obtained from two sources: an automatic electronic recording device and paper records. The automatic device had an inter...
time-series, multivariateFor complete information see the official challenge page: [Web Link]
clustering, sequential, causal-discovery, time-series, multivariate, domain-theoryPart of the problem in using an automated program to discover the unknown target function is to decide how to encode names such that the program can be ...
text, univariate, classificationTwo approaches were tested: a collaborative filter technique and a contextual approach. (i) The collaborative filter technique used only one file i.e...
multivariateThe Berkeley Multimodal Human Action Database (MHAD) contains 11 actions performed by 7 male and 5 female subjects in the range 23-30 years of age excep...
recognition, motion, action, classification, multiview52 columns for 52 weeks; normalised values of provided too.
time-series, multivariate, clusteringThe UMD Dynamic Scene Recognition dataset consists of 13 classes and 10 videos per class and is used to classify dynamic scenes. The dataset has been ...
video, motion, dynamic, classification, scene, recognitionParis-rue-Madame dataset contains 3D Mobile Laser Scanning (MLS) data from rue Madame, a street in the 6th Parisian district (France). The test zone con...
segmentation, 3d, semantic, classification, pointcloud, laserData was used to test 2 tier approach with learning from positive and negative examples
multivariateToday, we introduce Open Images, a dataset consisting of ~9 million URLs to images that have been annotated with labels spanning over 6000 categories. W...
annotation, deep, classification, real, large-scale, image, category, automaticThis data was collected from text ads found on twelve websites that deal with various farm animal related topics. Information from the ad creative and ...
text, classificationCSV format where each row is a paper and each column an attribute.
multivariate, clusteringThe data was retrieved from a set of 53500 CT images from 74 different patients (43 male, 31 female). Each CT slice is described by two histograms in p...
regression, domain-theoryThe Visual Attributes dataset contains visual attribute annotations for over 500 object classes (animate and inanimate) which are all represented in Ima...
object, recognition, attribute, classification, imagenetISPRS Test Project on Urban Classification and 3D Building Reconstruction The ISPRS working group III/4 announces the release of the 2D semantic label...
urban, reconstruction, recognition, building, 3d, classification, city, semanticFrom the original readme file (please consult it for more information): ------------------------- The documents in the Reuters-21578 collection appeared...
text, classificationThe Person Re-ID (PRID) 2011 dataset was created in co-operation with the Austrian Institute of Technology for the purpose of testing person re-identifi...
trajectory, graz, illumination, pedestrian, change, appearance, identification, classification, multiviewFor Each feature, a 64 element vector is given per sample of leaf. These vectors are taken as a contigous descriptors (for shape) or histograms (for tex...
classificationThis data comes from a survey conducted by the Graphics and Visualization Unit at Georgia Tech October 10 to November 16, 1997. The full details of the ...
multivariateThe data is in the transactional form. It contains the Latin names (species or genus) and state abbreviations.
multivariate, clusteringThis is a data set used by Ning Qian and Terry Sejnowski in their study using a neural net to predict the secondary structure of certain globular protei...
classification, sequentialThe Chars74K dataset consists of 64 classes (0-9, A-Z, a-z), 7705 characters obtained from natural images, 3410 hand drawn characters using a tablet PC,...
text recognition, text detection, classificationThe Oxford RobotCar Dataset contains over 100 repetitions of a consistent route through Oxford, UK, captured over a period of over a year. The dataset c...
driving, street, urban, time, recognition, autonomous, video, segmentation, robot, classification, detection, car, yearThe dataset collects data from an Android smartphone positioned in the chest pocket. Accelerometer Data are collected from 22 participants walking in th...
univariate, classification, clustering, sequential, time-seriesThe UBO 2014 consists of 7 semantic categories. Each of these 7 material categories contains measurements of 12 different material instances for being c...
illumination, material, classification, texture, light, recognitionThe CALTECH 256 dataset by Li Fei-Fei contains 30607 images for 256 categories.
object, detection, image, centered, classification, scene1. Protocol: Seven male and three female subjects (age 25 to 30), who have experienced aggression in scenarios such as physical fighting, took par...
time-series, classificationThe USAA dataset includes 8 different semantic class videos which are home videos of social occassions which feature activities of group of people. It c...
classificationOne of the challenges faced by our research was the unavailability of reliable training datasets. In fact this challenge faces any researcher in the fie...
classificationThe characters here were used for a PhD study on primitive extraction using HMM based models. The data consists of 2858 character samples, contained in ...
time-series, classification, clusteringThe Textures volume currently contains 154 images, all monochrome, 129 512x512 and 25 1024x1024. For the Brodatz texture images, the number in parenth...
segmentation, benchmark, evaluation, classification, synthetic, textureBelgiumTS is a large dataset with 10000+ traffic sign annotations, thousands of physically distinct traffic signs. 4 video sequences recorded with 8 hig...
urban, sign, belgium, road, traffic, classification, camera, calibrationMIL data sets used in our 2002 NIPS paper for Elepphant, Musk, TREC http://www.cs.cmu.edu/~juny/MILL/MIL-experiments.htm
classification, machine learningThe data has been produced using Monte Carlo simulations. The first 21 features (columns 2-22) are kinematic properties measured by the particle detecto...
classificationThis is the database of biological images (from the genetics model system, Drosophila melanogaster, a fruit fly) across multiple levels of variation. ...
classification, fly, biology, animal, variation, geneticMany different labeled video datasets have been collected over the past few years, but it is hard to compare them at a glance. So we have created a hand...
video, object, benchmark, classification, recognition, detection, action