BlogFeedback

Automobile

This data set consists of three types of entities: (a) the specification of an auto in terms of various characteristics, (b) its assigned insurance risk r…

multivariate, regression

Others

Energy efficiency

We perform energy analysis using 12 different building shapes simulated in Ecotect. The buildings differ with respect to the glazing area, the glazing are…

classification, multivariate, regression

Computer Science

Online Video Characteristics …

The presented dataset is composed of two tsv files named 'youtube_videos.tsv' and 'transcoding_mesurment.tsv'. The first contains 10 columns of fundament…

multivariate, regression

Computer Science

Communities and Crime

Many variables are included so that algorithms that select or learn weights for attributes could be tested. However, clearly unrelated attributes were n…

multivariate, regression

Social Sciences

Yacht Hydrodynamics

Prediction of residuary resistance of sailing yachts at the initial design stage is of a great value for evaluating the ships performance and for estimati…

multivariate, regression

Physical Systems

Stock portfolio performance

There are three disadvantages of weighted scoring stock selection models. First, they cannot identify the relations between weights of stock-picking conce…

multivariate, regression

Business

Twin gas sensor arrays

This dataset includes the recordings of five replicates of an 8-sensor array. Each unit holds 8 MOX sensors and integrates custom-designed electronics for…

classification, domain-theory, multivariate, regression, time-series

Computer Science

NoisyOffice

AIMS AND PURPOSES This corpus is intended to do cleaning (or binarization) and enhancement of noisy grayscale printed text images using supervised learni…

classification, multivariate, regression

Computer Science

SkillCraft1 Master Table Data…

-- We aggregated screen movements into screen-fixations using a Salvucci & Goldberg (2000) dispersion-threshold algorithm, and defined Perception Actio…

multivariate, regression

Games

Beijing PM2.5 Data

The datas time period is between Jan 1st, 2010 to Dec 31st, 2014. Missing data are denoted as NA.

multivariate, regression, time-series

Physical Systems

Tamilnadu Electricity Board H…

Collect the real time readings for residential,commercial,industrial,agriculure,to find the accuracy consumption in Tamil Nadu Around Thanajvur

classification, clustering, multivariate, regression

Life Sciences

KEGG Metabolic Relation Netwo…

KEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. In …

classification, clustering, multivariate, regression, text, univariate

Life Sciences

YearPredictionMSD

You should respect the following train / test split: train: first 463,715 examples test: last 51,630 examples It avoids the 'producer effect' by making su…

multivariate, regression

Others

Gas Sensor Array Drift Datase…

This data set contains 13,910 measurements from 16 chemical sensors exposed to 6 gases at different concentration levels. This dataset is an extension of …

causa, classification, clustering, multivariate, regression, time-series

Computer Science

Individual household electric…

This archive contains 2075259 measurements gathered between December 2006 and November 2010 (47 months). Notes: 1.(global_active_power*1000/60 - sub_mete…

clustering, multivariate, regression, time-series

Physical Systems

Facebook metrics

The data is related to posts' published during the year of 2014 on the Facebook's page of a renowned cosmetics brand. This dataset contains 500 of the 790…

multivariate, regression

Business

Concrete Slump Test

The data set includes 103 data points. There are 7 input variables, and 3 output variables in the data set. The initial data set included 78 data. After s…

multivariate, regression

Computer Science

PM2.5 Data of Five Chinese Ci…

The time period is between Jan 1st, 2010 to Dec 31st, 2015. Missing data are denoted as NA.

multivariate, regression, time-series

Physical Systems

DrivFace

The DrivFace database contains images sequences of subjects while driving in real scenarios. It is composed of 606 samples of 640480 pixels each, acquired…

classification, clustering, multivariate, regression

Computer Science

Air Quality

The dataset contains 9358 instances of hourly averaged responses from an array of 5 metal oxide chemical sensors embedded in an Air Quality Chemical Multi…

multivariate, regression, time-series

Computer Science

GPS Trajectories

The dataset is composed by two tables. The first table go_track_tracks presents general attributes and each instance has one trajectory that is represente…

classification, multivariate, regression

Computer Science

Greenhouse Gas Observing Netw…

This data set contains time series of greenhouse gas (GHG) concentrations at 2921 grid cells in California created using simulations of the Weather Resear…

multivariate, regression, time-series

Physical Systems

Computer Hardware

The estimated relative performance values were estimated by the authors using a linear regression method. See their article (pp 308-313) for more details…

multivariate, regression

Computer Science

Gas sensor array exposed to t…

A chemical detection platform composed of 8 chemo-resistive gas sensors was exposed to turbulent gas mixtures generated naturally in a wind tunnel. The ac…

classification, multivariate, regression, time-series

Computer Science

Forest Fires

In [Cortez and Morais, 2007], the output 'area' was first transformed with a ln(x+1) function. Then, several Data Mining methods were applied. After fi…

multivariate, regression

Physical Systems

Geographical Original of Music

The dataset was built from a personal collection of 1059 tracks covering 33 countries/area. The music used is traditional, ethnic or `world' only, as cla…

classification, multivariate, regression

Others

KEGG Metabolic Reaction Netwo…

KEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. In …

classification, clustering, multivariate, regression, text, univariate

Life Sciences

Facebook Comment Volume Datas…

The Dataset is uploaded in ZIP format. The dataset contains 5 variants of the dataset, for the details about the variants and detailed analysis read and c…

multivariate, regression

Others

Challenger USA Space Shuttle …

There are two databases: (both use the same set of 5 attributes): 1. Primary o-ring erosion and/or blowby 2. Primary o-ring erosion only The two database…

multivariate, regression

Physical Systems

Open University Learning Anal…

Open University Learning Analytics Dataset (OULAD) contains data about courses, students and their interactions with Virtual Learning Environment (VLE) fo…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

Auto MPG

This dataset is a slightly modified version of the dataset provided in the StatLib library. In line with the use by Ross Quinlan (1993) in predicting the…

multivariate, regression

Others

Gas sensor array under dynami…

This data set contains the acquired time series from 16 chemical sensors exposed to gas mixtures at varying concentration levels. In particular, we genera…

classification, multivariate, regression, time-series

Computer Science

Concrete Compressive Strength

Number of instances 1030 Number of Attributes 9 Attribute breakdown 8 quantitative input variables, and 1 quantitative output variable Missing Attribute …

multivariate, regression

Physical Systems

Insurance Company Benchmark (…

Information about customers consists of 86 variables and includes product usage data and socio-demographic data derived from zip area codes. The data was …

description, multivariate, regression

Social Sciences

Cargo 2000 Freight Tracking a…

A description of the underlying Cargo 2000 standard and the processes reflected in the data set can be found at [Web Link].

classification, multivariate, regression, sequential

Business

SML2010

The dataset could contain missing values. The data was sampled every minute, computing and uploading it smoothed with 15 minute means. The header of the d…

multivariate, regression, sequential, text, time-series

Computer Science

Breast Cancer Wisconsin (Prog…

Each record represents follow-up data for one breast cancer case. These are consecutive patients seen by Dr. Wolberg since 1984, and include only those c…

classification, multivariate, regression

Life Sciences

Airfoil Self-Noise

The NASA data set comprises different size NACA 0012 airfoils at various wind tunnel speeds and angles of attack. The span of the airfoil and the observer…

multivariate, regression

Physical Systems

Educational Process Mining (E…

The experiments have been carried out with a group of 115 students of first-year, undergraduate Engineering major of the University of Genoa. We carried…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

Condition Based Maintenance o…

The experiments have been carried out by means of a numerical simulator of a naval vessel (Frigate) characterized by a Gas Turbine (GT) propulsion plant. …

multivariate, regression

Computer Science

Fertility

Provide all relevant information about your data set.

classification, multivariate, regression

Life Sciences

Appliances energy prediction

The data set is at 10 min for about 4.5 months. The house temperature and humidity conditions were monitored with a ZigBee wireless sensor network. Each w…

multivariate, regression, time-series

Computer Science

Wine Quality

The two datasets are related to red and white variants of the Portuguese "Vinho Verde" wine. For more details, consult: [Web Link] or the reference [Corte…

classification, multivariate, regression

Business

Online News Popularity

* The articles were published by Mashable (www.mashable.com) and their content as the rights to reproduce it belongs to them. Hence, this dataset does not…

classification, multivariate, regression

Business

UJIIndoorLoc

Many real world applications need to know the localization of a user in the world to provide their services. Therefore, automatic user localization has be…

classification, multivariate, regression

Computer Science

Student Performance

This data approach student achievement in secondary education of two Portuguese schools. The data attributes include student grades, demographic, social a…

classification, multivariate, regression

Social Sciences

UJIIndoorLoc-Mag

Indoor localization is a key topic for mobile computing. However, it is still very difficult for the mobile sensing community to compare state-of-art Indo…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

Parkinson Disease Spiral Draw…

The PD and control handwriting database consists of 62 PWP (People with parkinson) and 15 healthy individuals who appealed at the Department of Neurology …

classification, clustering, multivariate, regression

Computer Science

Tennis Major Tournament Match…

N/A

classification, clustering, multivariate, regression

Others

Servo

Ross Quinlan: This data was given to me by Karl Ulrich at MIT in 1986. I didn't record his description at the time, but here's his subsequent (1992) rec…

multivariate, regression

Computer Science

Geo-Magnetic field and WLAN d…

Indoor localisation is a key topic for the Ambient Intelligence (AmI) research community. In this scenarios, recent advancements in wearable technologie…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

KDD Cup 1998 Data

Please see associated text files in the download folder.

multivariate, regression

Others

Gas sensor array under flow m…

The measured data was collected using a chemical sensing system based on an array of 16 metal-oxide gas sensors and an external mechanical ventilator to s…

classification, multivariate, regression, time-series

Computer Science

KDC-4007 dataset Collection

The most important feature of this dataset is its simplicity to use and its being well-documented, which can be widely used in various studies of text ana…

classification, multivariate, regression, text

Computer Science

wiki4HE

Ongoing research on university faculty perceptions and practices of using Wikipedia as a teaching resource. Based on a Technology Acceptance Model, the re…

causal-discovery, clustering, multivariate, regression

Social Sciences

Parkinson Speech Dataset with…

The PD database consists of training and test files. The training data belongs to 20 PWP (6 female, 14 male) and 20 healthy individuals (10 female, 10 mal…

classification, multivariate, regression

Life Sciences

Communities and Crime Unnorma…

The source datasets needed to be combined via programming. Many variables are included so that algorithms that select or learn weights for attributes coul…

multivariate, regression

Social Sciences

ISTANBUL STOCK EXCHANGE

Data is collected from imkb.gov.tr and finance.yahoo.com. Data is organized with regard to working days in Istanbul Stock Exchange.

classification, multivariate, regression, time-series, univariate

Business

Parkinsons Telemonitoring

This dataset is composed of a range of biomedical voice measurements from 42 people with early-stage Parkinson's disease recruited to a six-month trial of…

multivariate, regression

Life Sciences

Cuff-Less Blood Pressure Esti…

The main goal of this data set is providing clean and valid signals for designing cuff-less blood pressure estimation algorithms. The raw electrocardiogra…

classification, multivariate, regression

Life Sciences

Buzz in social media

Please see [Web Link]

classification, multivariate, regression, time-series

Computer Science

Solar Flare

Notes: -- The database contains 3 potential classes, one for the number of times a certain type of solar flare occured in a 24 hour period. -- Each…

multivariate, regression

Physical Systems

Physicochemical Properties of…

Provide all relevant information about your data set.

multivariate, regression

Life Sciences

Combined Cycle Power Plant

The dataset contains 9568 data points collected from a Combined Cycle Power Plant over 6 years (2006-2011), when the power plant was set to work with full…

multivariate, regression

Computer Science

Robot Execution Failures

The donation includes 5 datasets, each of them defining a different learning problem: * LP1: failures in approach to grasp position * LP2: failur…

classification, multivariate, time-series

Physical Systems

Anuran Calls (MFCCs)

This dataset was used in several classifications tasks related to the challenge of anuran species recognition through their calls. It is a multilabel data…

classification, clustering, multivariate

Life Sciences

seismic-bumps

Mining activity was and is always connected with the occurrence of dangers which are commonly called mining hazards. A special case of such threat is a s…

classification, multivariate

Others

Polish companies bankruptcy d…

The dataset is about bankruptcy prediction of Polish companies. The data was collected from Emerging Markets Information Service (EMIS, [Web Link]), which…

classification, multivariate

Business

SECOM

A complex modern semi-conductor manufacturing process is normally under consistent surveillance via the monitoring of signals/variables collected from sen…

causal-discovery, classification, multivariate

Computer Science

KDD Cup 1999 Data

Please see task description.

classification, multivariate

Computer Science

SIFT10M

In SIFT10M, the titles of the png files indicate the columns position of the SIFT features. This data set has been used for evaluating the approximate nea…

causal-discovery, multivariate

Computer Science

Heart Disease

This database contains 76 attributes, but all published experiments refer to using a subset of 14 of them. In particular, the Cleveland database is the o…

classification, multivariate

Life Sciences

Wine

These data are the results of a chemical analysis of wines grown in the same region in Italy but derived from three different cultivars. The analysis dete…

classification, multivariate

Physical Systems

Labor Relations

Data was used to test 2 tier approach with learning from positive and negative examples

multivariate

Social Sciences

User Knowledge Modeling

-- The users' knowledge class were classified by the authors using intuitive knowledge classifier (a hybrid ML technique of k-NN and meta-heuristic…

classification, clustering, multivariate

Computer Science

Echocardiogram

All the patients suffered heart attacks at some point in the past. Some are still alive and some are not. The survival and still-alive variables, when ta…

classification, multivariate

Life Sciences

Letter Recognition

The objective is to identify each of a large number of black-and-white rectangular pixel displays as one of the 26 capital letters in the English alphabet…

classification, multivariate

Computer Science

Madelon

MADELON is an artificial dataset containing data points grouped in 32 clusters placed on the vertices of a five dimensional hypercube and randomly labeled…

classification, multivariate

Others

Haberman's Survival

The dataset contains cases from a study that was conducted between 1958 and 1970 at the University of Chicago's Billings Hospital on the survival of patie…

classification, multivariate

Life Sciences

KASANDR

We created this data by sampling and processing the www.kelkoo.com logs. The data records offers which were clicked (or shown) to the users of the www.kel…

causal-discovery, multivariate

Life Sciences

Breast Tissue

Impedance measurements were made at the frequencies: 15.625, 31.25, 62.5, 125, 250, 500, 1000 KHz Impedance measurements of freshly excised breast tissue …

classification, multivariate

Life Sciences

Waveform Database Generator (…

Notes: -- 3 classes of waves -- 40 attributes, all of which include noise -- The latter 19 attributes are all noise attributes with mean…

classification, generator, multivariate

Physical Systems

Abscisic Acid Signaling Netwo…

The objective is to determine the set of boolean rules that describe the interactions of the nodes within this plant signaling network. The dataset includ…

causal-discovery, multivariate

Life Sciences

Statlog (Australian Credit Ap…

This file concerns credit card applications. All attribute names and values have been changed to meaningless symbols to protect confidentiality of the da…

classification, multivariate

Financial

Grammatical Facial Expressions

The automated analysis of facial expressions has been widely used in different research areas, such as biometrics or emotional analysis. Special import…

classification, clustering, multivariate, sequential

Computer Science

Daily and Sports Activities

Brief Description of the Dataset: --------------------------------- Each of the 19 activities is performed by eight subjects (4 female, 4 male, between th…

classification, clustering, multivariate, time-series

Computer Science

Plants

The data is in the transactional form. It contains the Latin names (species or genus) and state abbreviations.

clustering, multivariate

Life Sciences

Gastrointestinal Lesions in R…

This dataset contains the features extracted from a database of colonoscopic videos showing gastrointestinal lesions. It also contains the ground truth co…

classification, multivariate

Computer Science

seeds

The examined group comprised kernels belonging to three different varieties of wheat: Kama, Rosa and Canadian, 70 elements each, randomly selected for the…

classification, clustering, multivariate

Life Sciences

Post-Operative Patient

The classification task of this database is to determine where patients in a postoperative recovery area should be sent to next. Because hypothermia is a…

classification, multivariate

Life Sciences

Pen-Based Recognition of Hand…

We create a digit database by collecting 250 samples from 44 writers. The samples written by 30 writers are used for training, cross-validation and writer…

classification, multivariate

Computer Science

Record Linkage Comparison Pat…

The records represent individual data including first and family name, sex, date of birth and postal code, which were collected through iterative insertio…

classification, multivariate

Others

Chess (King-Rook vs. King)

An Inductive Logic Programming (ILP) or relational learning framework is assumed (Muggleton, 1992). The learning system is provided with examples of chess…

classification, multivariate

Games

MoCap Hand Postures

A Vicon motion capture camera system was used to record 12 users performing 5 hand postures with markers attached to a left-handed glove. A rigid patter…

classification, clustering, multivariate

Computer Science

Multiple Features

This dataset consists of features of handwritten numerals (`0'--`9') extracted from a collection of Dutch utility maps. 200 patterns per class (for a tota…

classification, multivariate

Computer Science

Soybean (Large)

There are 19 classes, only the first 15 of which have been used in prior work. The folklore seems to be that the last four classes are unjustified by the …

classification, multivariate

Life Sciences

Internet Usage Data

This data comes from a survey conducted by the Graphics and Visualization Unit at Georgia Tech October 10 to November 16, 1997. The full details of the su…

multivariate

Computer Science

Census-Income (KDD)

This data set contains weighted census data extracted from the 1994 and 1995 Current Population Surveys conducted by the U.S. Census Bureau. The data cont…

classification, multivariate

Social Sciences

Arrhythmia

This database contains 279 attributes, 206 of which are linear valued and the rest are nominal. Concerning the study of H. Altay Guvenir: "The aim is to…

classification, multivariate

Life Sciences

LED Display Domain

This simple domain contains 7 Boolean attributes and 10 concepts, the set of decimal digits. Recall that LED displays contain 7 light-emitting diodes -- …

classification, generator, multivariate

Computer Science

Low Resolution Spectrometer

The Infra-Red Astronomy Satellite (IRAS) was the first attempt to map the full sky at infra-red wavelengths. This could not be done from ground observato…

classification, multivariate

Physical Systems

Artificial Characters

This database has been artificially generated by using a first order theory which describes the structure of ten capital letters of the English alphabet a…

classification, multivariate

Computer Science

PAMAP2 Physical Activity Moni…

The PAMAP2 Physical Activity Monitoring dataset contains data of 18 different physical activities (such as walking, cycling, playing soccer, etc.), perfor…

classification, multivariate, time-series

Computer Science

StoneFlakes

Background information: The data set concerns the earliest history of mankind. Prehistoric men created the desired shape of a stone tool by striking on a …

causal-discovery, classification, clustering, multivariate

Others

Page Blocks Classification

The 5473 examples comes from 54 distinct documents. Each observation concerns one block. All attributes are numeric. Data are in a format readable by C4.5.

classification, multivariate

Computer Science

University

Format: Each observation concerns one university. In some cases, more information is provided about the attribute (e.g., units or domain). Some duplicates…

classification, multivariate

Others

Movie

The data is stored in relational form across several files. The central file (MAIN) is a list of movies, each with a unique identifier. These identifiers …

multivariate, relational

Others

Spambase

The "spam" concept is diverse: advertisements for products/web sites, make money fast schemes, chain letters, pornography... Our collection of spam e-mai…

classification, multivariate

Computer Science

EEG Database

This data arises from a large study to examine EEG correlates of genetic predisposition to alcoholism. It contains measurements from 64 electrodes placed …

multivariate, time-series

Life Sciences

Reuter_50_50

The dataset is the subset of RCV1. These corpus has already been used in author identification experiments. In the top 50 authors (with respect to total s…

classification, clustering, domain-theory, multivariate, text

Computer Science

Leaf

For further details on this dataset and/or its attributes, please read the 'ReadMe.pdf' file included and/or consult the Master's Thesis 'Development of a…

classification, multivariate

Computer Science

Audiology (Standardized)

This database is a standardized version of the original audiology database (see audiology.* in this directory). The non-standard set of attributes have b…

classification, multivariate

Life Sciences

Urban Land Cover

Contains training and testing data for classifying a high resolution aerial image into 9 types of urban land cover. Multi-scale spectral, size, shape, and…

classification, multivariate

Physical Systems

REALDISP Activity Recognition…

The REALDISP (REAListic sensor DISPlacement) dataset has been originally collected to investigate the effects of sensor displacement in the activity recog…

classification, multivariate, time-series

Computer Science

Connect-4

This database contains all legal 8-ply positions in the game of connect-4 in which neither player has won yet, and in which the next move is not forced. …

classification, multivariate, spatial

Games

Pittsburgh Bridges

There are two versions to the database: - V1 contains the original examples and - V2 contains descriptions after discretizing numeric properti…

classification, multivariate

Others

Connectionist Bench (Nettalk …

This is an updated and corrected version of the data set used by Sejnowski and Rosenberg in their influential study of speech generation using a neural ne…

multivariate

Others

Lung Cancer

This data was used by Hong and Young to illustrate the power of the optimal discriminant plane even in ill-posed settings. Applying the KNN method in the …

classification, multivariate

Life Sciences

Online Retail

This is a transnational data set which contains all the transactions occurring between 01/12/2010 and 09/12/2011 for a UK-based and registered non-store o…

classification, clustering, multivariate, sequential, time-series

Business

Quality Assessment of Digital…

* The dataset was acquired and annotated by professional physicians at 'Hospital Universitario de Caracas'. * The subjective judgments (target variables) …

classification, multivariate

Life Sciences

Corel Image Features

The original image collection was obtained from Corel at [Web Link]. There are 68,040 photo images from various categories. Each set of features is store…

multivariate

Others

Annealing

N/A

classification, multivariate

Physical Systems

Dota2 Games Results

Dota 2 is a popular computer game with two teams of 5 players. At the start of the game each player chooses a unique hero with different strengths and wea…

classification, multivariate

Games

Occupancy Detection

Three data sets are submitted, for training and testing. Ground-truth occupancy was obtained from time stamped pictures that were taken every minute. For …

classification, multivariate, time-series

Computer Science

YouTube Multiview Video Games…

Please see the README for the details on the data organization, and so on.

classification, clustering, multivariate, text

Computer Science

Lymphography

This is one of three domains provided by the Oncology Institute that has repeatedly appeared in the machine learning literature. (See also breast-cancer a…

classification, multivariate

Life Sciences

Pima Indians Diabetes

Several constraints were placed on the selection of these instances from a larger database. In particular, all patients here are females at least 21 year…

classification, multivariate

Life Sciences

Mushroom

This data set includes descriptions of hypothetical samples corresponding to 23 species of gilled mushrooms in the Agaricus and Lepiota Family (pp. 500-52…

classification, multivariate

Life Sciences

Cloud

The data sets we propose to analyse are constituted of 1024 vectors, each vector includes 10 parameters. You can think of it as a 1024*10 matrix. To produ…

multivariate

Physical Systems

Qualitative_Bankruptcy

The parameters which we used for collecting the dataset is referred from the paper 'The discovery of experts decision rules from qualitative bankruptcy da…

classification, multivariate

Computer Science

EEG Eye State

All data is from one continuous EEG measurement with the Emotiv EEG Neuroheadset. The duration of the measurement was 117 seconds. The eye state was detec…

classification, multivariate, sequential, time-series

Life Sciences

Water Treatment Plant

This dataset comes from the daily measures of sensors in a urban waste water treatment plant. The objective is to classify the operational state of the pl…

clustering, multivariate

Physical Systems

Mammographic Mass

Mammography is the most effective method for breast cancer screening available today. However, the low positive predictive value of breast biopsy resultin…

classification, multivariate

Life Sciences

Hayes-Roth

This database contains 5 numeric-valued attributes. Only a subset of 3 are used during testing (the latter 3). Furthermore, only 2 of the 3 concepts are…

classification, multivariate

Social Sciences

Northix

Northix is designed to be a schema matching benchmark problem for data integration of two entity relationship databases. Northix is the resulting schema m…

classification, multivariate, text, univariate

Computer Science

Contraceptive Method Choice

This dataset is a subset of the 1987 National Indonesia Contraceptive Prevalence Survey. The samples are married women who were either not pregnant or do …

classification, multivariate

Life Sciences

Statlog (Shuttle)

Approximately 80% of the data belongs to class 1. Therefore the default accuracy is about 80%. The aim here is to obtain an accuracy of 99 - 99.9%. The e…

classification, multivariate

Physical Systems

AAAI 2013 Accepted Papers

CSV format where each row is a paper and each column is an attribute.

clustering, multivariate

Computer Science

Motion Capture Hand Postures

A Vicon motion capture camera system was used to record 12 users performing 5 hand postures with markers attached to a left-handed glove. A rigid patter…

classification, clustering, multivariate

Computer Science

MHEALTH Dataset

The MHEALTH (Mobile HEALTH) dataset comprises body motion and vital signs recordings for ten volunteers of diverse profile while performing several physic…

classification, multivariate, time-series

Computer Science

Covertype

Predicting forest cover type from cartographic variables only (no remotely sensed data). The actual forest cover type for a given observation (30 x 30 me…

classification, multivariate

Life Sciences

Las Vegas Strip

All the 504 reviews were collected between January and August of 2015.

classification, regression

Business

Predict keywords activities i…

See files and/or [Web Link]

multivariate, sequential, time-series

Computer Science

Balance Scale

This data set was generated to model psychological experimental results. Each example is classified as having the balance scale tip to the right, tip to …

classification, multivariate

Social Sciences

Statlog (Heart)

Cost Matrix _______ abse pres absence 0 1 presence 5 0 where the rows represent the true values and the columns the predicted.

classification, multivariate

Life Sciences

ElectricityLoadDiagrams201120…

Data set has no missing values. Values are in kW of each 15 min. To convert values in kWh values must be divided by 4. Each column represent one client. S…

clustering, regression, time-series

Computer Science

microblogPCU

Our dataset is used by us to explore spammers in microblog and you can access our demo system at [Web Link]Please add :8080 after the domain name as port.…

causal-discovery, classification, multivariate, sequential, text, univariate

Computer Science

Crowdsourced Mapping

This dataset was derived from geospatial data from two sources: 1) Landsat time-series satellite imagery from the years 2014-2015, and 2) crowdsourced geo…

classification, multivariate

Physical Systems

Connectionist Bench (Sonar, M…

The file "sonar.mines" contains 111 patterns obtained by bouncing sonar signals off a metal cylinder at various angles and under various conditions. The …

classification, multivariate

Physical Systems

Statlog (Vehicle Silhouettes)

The purpose is to classify a given silhouette as one of four types of vehicle, using a set of features extracted from the silhouette. The vehicle may be …

classification, multivariate

Others

ILPD (Indian Liver Patient Da…

This data set contains 416 liver patient records and 167 non liver patient records.The data set was collected from north east of Andhra Pradesh, India. Se…

classification, multivariate

Life Sciences

Sponge

These are atlantic-mediterranean marine sponges that belong to O.Hadromerida (Demospongiae.Porifera).

clustering, multivariate

Life Sciences

Spoken Arabic Digit

Dataset from 8800(10 digits x 10 repetitions x 88 speakers) time series of 13 Frequency Cepstral Coefficients (MFCCs) had taken from 44 males and 44 femal…

classification, multivariate, time-series

Others

Australian Sign Language sign…

Data was captured using a setup that consisted of: - Two Fifth Dimension Technologies (5DT) gloves, one right and one left - Two Ascension Flock-of-Bir…

classification, multivariate, time-series

Others

MicroMass

This MALDI-TOF dataset consists in:A) A reference panel of 20 Gram positive and negative bacterial species covering 9 genera among which several species a…

classification, multivariate

Life Sciences

Flags

This data file contains details of various nations and their flags. In this file the fields are separated by spaces (not commas). With this data you can …

classification, multivariate

Others

p53 Mutants

Biophysical models of mutant p53 proteins yield features which can be used to predict p53 transcriptional activity. All class labels are determined via i…

classification, multivariate

Life Sciences

MONK's Problems

The MONK's problem were the basis of a first international comparison of learning algorithms. The result of this comparison is summarized in "The MONK's P…

classification, multivariate

Others

Restaurant & consumer data

Two approaches were tested: a collaborative filter technique and a contextual approach. (i) The collaborative filter technique used only one file i.e.,…

multivariate

Computer Science

Forest type mapping

This data set contains training and testing data from a remote sensing study which mapped different forest types based on their spectral characteristics a…

classification, multivariate

Life Sciences

Statlog (Image Segmentation)

The instances were drawn randomly from a database of 7 outdoor images. The images were handsegmented to create a classification for every pixel. Each …

classification, multivariate

Others

URL Reputation

Uncompressing the archive url_svmlight.tar.gz will yield a directory url_svmlight/ containing the following files: * FeatureTypes --- A text file list…

classification, multivariate, time-series

Computer Science

First-order theorem proving

See the file bridge-holden-paulson-details.txt in the submitted tarball.

classification, multivariate

Computer Science

MAGIC Gamma Telescope

The data are MC generated (see below) to simulate registration of high energy gamma particles in a ground-based atmospheric Cherenkov gamma telescope usin…

classification, multivariate

Physical Systems

Gas Sensor Array Drift Dataset

This archive contains 13910 measurements from 16 chemical sensors utilized in simulations for drift compensation in a discrimination task of 6 gases at va…

classification, multivariate

Computer Science

Bike Sharing Dataset

Bike sharing systems are new generation of traditional bike rentals where whole process from membership, rental and return back has become automatic. Thro…

regression, univariate

Social Sciences

MiniBooNE particle identifica…

The submitted file is set up as follows. In the first line is the the number of signal events followed by the number of background events. The signal even…

classification, multivariate

Physical Systems

Diabetes 130-US hospitals for…

The dataset represents 10 years (1999-2008) of clinical care at 130 US hospitals and integrated delivery networks. It includes over 50 features representi…

classification, clustering, multivariate

Life Sciences

LSVT Voice Rehabilitation

The original paper demonstrated that it is possible to correctly replicate the experts' binary assessment with approximately 90% accuracy using both 10-fo…

classification, multivariate

Life Sciences

Statlog (Landsat Satellite)

The database consists of the multi-spectral values of pixels in 3x3 neighbourhoods in a satellite image, and the classification associated with the centra…

classification, multivariate

Physical Systems

Climate Model Simulation Cras…

This dataset contains records of simulation crashes encountered during climate model uncertainty quantification (UQ) ensembles. Ensemble members were co…

classification, multivariate

Physical Systems

Coil 1999 Competition Data

This data comes from a water quality study where samples were taken from sites on different European rivers of a period of approximately one year. These s…

multivariate

Physical Systems

Dataset for Sensorless Drive …

Features are extracted from electric current drive signals. The drive has intact and defective components. This results in 11 different classes with diffe…

classification, multivariate

Computer Science

Australian Sign Language signs

The source of the data is the raw measurements from a Nintendo PowerGlove. It was interfaced through a PowerGlove Serial Interface to a Silicon Graphics 4…

classification, multivariate, time-series

Others

UJI Pen Characters

We create a character database by collecting samples from 11 writers. Each writer contributed with letters (lower and uppercase), digits, and other chara…

classification, multivariate, sequential

Computer Science

Dataset for ADL Recognition w…

The Dataset for ADL Recognition with Wrist-worn Accelerometer is a public collection of labelled accelerometer data recordings to be used for the creation…

classification, clustering, multivariate, time-series

Computer Science

PubChem Bioassay Data

21 bioassay datasets generated from Pubchem. Both Primary and confirmatory bioassays (12 bioassays, 21 mixes)The data is provided in the same train/test s…

classification, multivariate

Life Sciences

Poker Hand

Each record is an example of a hand consisting of five playing cards drawn from a standard deck of 52. Each card is described using two attributes (suit a…

classification, multivariate

Games

Chess (King-Rook vs. King-Kni…

The companion file is a Common Lisp demonstration file that generates knight-pin Chess end-game samples. Start up Lisp and load the file. It generates 10…

classification, generator, multivariate

Games

Car Evaluation

Car Evaluation Database was derived from a simple hierarchical decision model originally developed for the demonstration of DEX, M. Bohanec, V. Rajkovic: …

classification, multivariate

Others

Wall-Following Robot Navigati…

The provided files comprise three different data sets. The first one contains the raw values of the measurements of all 24 ultrasound sensors and the cor…

classification, multivariate, sequential

Computer Science

Hybrid Indoor Positioning Dat…

The measurements were created to ease the development, comparison and evaluation of fingerprinting based hybrid indoor positioning methods. The measuremen…

classification, multivariate, sequential, time-series

Computer Science

Amazon Commerce reviews set

dataset are derived from the customers reviews in Amazon Commerce Website for authorship identification. Most previous studies conducted the identifi…

classification, domain-theory, multivariate, text

Physical Systems

ser Knowledge Modeling Data (…

-- The users' knowledge class were classified by the authors using intuitive knowledge classifier (a hybrid ML technique of k-NN and meta-heuristic …

classification, multivariate

Computer Science

Diabetes

Diabetes patient records were obtained from two sources: an automatic electronic recording device and paper records. The automatic device had an interna…

multivariate, time-series

Life Sciences

UbiqLog (smartphone lifeloggi…

This is the first smartphone based lifelogging dataset that is going to be available for public use. Please consider that the user of this dataset are obl…

causal-discovery, multivariate

Computer Science

Demospongiae

This dataset contains 503 sponges belonging to the Demospongiae class collected from the Mediterranean (451 sponges) and Atlantic oceans (52 sponges). Eac…

classification, multivariate

Life Sciences

Chronic_Kidney_Disease

We use the following representation to collect the dataset age - age bp - blood pressure sg - specific gravity al - …

classification, multivariate

Others

Japanese Vowels

The data was collected for examining our newly developed classifier for multidimensional curves (multidimensional time series). Nine male speakers uttered…

classification, multivariate, time-series

Others

Soybean (Small)

A small subset of the original soybean database. See the reference for Fisher and Schlimmer in soybean-large.names for more information. Steven Souders …

classification, multivariate

Life Sciences

Thoracic Surgery Data

The data was collected retrospectively at Wroclaw Thoracic Surgery Centre for patients who underwent major lung resections for primary lung cancer in the …

classification, multivariate

Life Sciences

Musk (Version 1)

This dataset describes a set of 92 molecules of which 47 are judged by human experts to be musks and the remaining 45 molecules are judged to be non-musks…

classification, multivariate

Physical Systems

default of credit card clients

This research aimed at the case of customers default payments in Taiwan and compares the predictive accuracy of probability of default among six data mini…

classification, multivariate

Business

Libras Movement

The dataset (movement_libras) contains 15 classes of 24 instances each, where each class references to a hand movement type in LIBRAS. In the video pre-p…

classification, clustering, multivariate, sequential

Others

Vertebral Column

Biomedical data set built by Dr. Henrique da Mota during a medical residence period in the Group of Applied Research in Orthopaedics (GARO) of the Centre …

classification, multivariate

Others

Congressional Voting Records

This data set includes votes for each of the U.S. House of Representatives Congressmen on the 16 key votes identified by the CQA. The CQA lists nine diff…

classification, multivariate

Social Sciences

Diabetic Retinopathy Debrecen…

This dataset contains features extracted from the Messidor image set to predict whether an image contains signs of diabetic retinopathy or not. All featur…

classification, multivariate

Life Sciences

News Aggregator

News are grouped into clusters that represent pages discussing the same news story. The dataset includes also references to web pages that, at the access…

classification, clustering, multivariate

Others

Drug consumption (quantified)

Database contains records for 1885 respondents. For each respondent 12 attributes are known: Personality measurements which include NEO-FFI-R (neuroticism…

classification, multivariate

Social Sciences

TV News Channel Commercial De…

Automatic identification of commercial blocks in news videos finds a lot of applications in the domain of television broadcast analysis and monitoring. Co…

classification, clustering, multivariate

Computer Science

Parkinsons

This dataset is composed of a range of biomedical voice measurements from 31 people, 23 with Parkinson's disease (PD). Each column in the table is a parti…

classification, multivariate

Life Sciences

Smartphone-Based Recognition …

The experiments were carried out with a group of 30 volunteers within an age bracket of 19-48 years. They performed a protocol of activities composed of s…

classification, multivariate, time-series

Life Sciences

Taxi Service Trajectory - Pre…

For complete information see the official challenge page: [Web Link]

causal-discovery, clustering, domain-theory, multivariate, sequential, time-series

Computer Science

Semeion Handwritten Digit

1593 handwritten digits from around 80 persons were scanned, stretched in a rectangular box 16x16 in a gray scale of 256 values.Then each pixel of each im…

classification, multivariate

Computer Science

Ecoli

The references below describe a predecessor to this dataset and its development. They also give results (not cross-validated) for classification by a rule…

classification, multivariate

Life Sciences

Mechanical Analysis

F. Bergadano supplied this database. Each instance contains many components, each of which has 8 attributes. Different instances in this database have d…

classification, multivariate

Computer Science

Epileptic Seizure Recognition

Please find the original data at '[Web Link]'

classification, clustering, multivariate, time-series

Life Sciences

CNAE-9

This is a data set containing 1080 documents of free text business descriptions of Brazilian companies categorized into a subset of 9 categories cataloged…

classification, multivariate, text

Business

Dexter

The original data were formatted by Thorsten Joachims in the bag-of-words representation. There were 9947 features (of which 2562 are always zeros for all…

classification, multivariate

Others

Statlog (German Credit Data)

Two datasets are provided. the original dataset, in the form provided by Prof. Hofmann, contains categorical/symbolic attributes and is in the file "germ…

classification, multivariate

Financial

Cardiotocography

2126 fetal cardiotocograms (CTGs) were automatically processed and the respective diagnostic features measured. The CTGs were also classified by three exp…

classification, multivariate

Life Sciences

Human Activity Recognition Us…

The experiments have been carried out with a group of 30 volunteers within an age bracket of 19-48 years. Each person performed six activities (WALKING, W…

classification, clustering, multivariate, time-series

Computer Science

Indoor User Movement Predicti…

This dataset represents a real-life benchmark in the area of Ambient Assisted Living applications, as described in [1]. The binary classification task con…

classification, multivariate, sequential, time-series

Computer Science

NYSK

Documents are first obtained via a Web search using AMIEI: an integrated platform for delivering enterprise intelligence, developed by AMI Software ([Web …

clustering, multivariate, sequential, text

Social Sciences

MEU-Mobile KSD

The dataset is used in the evaluation of EER, FRR and FAR metrics using a new anomaly detector model (Med-Min-Diff). The typed text in the experiment is t…

classification, multivariate

Computer Science

Credit Approval

This file concerns credit card applications. All attribute names and values have been changed to meaningless symbols to protect confidentiality of the da…

classification, multivariate

Financial

Japanese Credit Screening

Examples represent positive and negative instances of people who were and were not granted credit. The theory was generated by talking to the individual…

classification, domain-theory, multivariate

Financial

Hepatitis

Please ask Gail Gong for further information on this database.

classification, multivariate

Life Sciences

Cylinder Bands

Here's the abstract from the above reference: ABSTRACT: Machine learning tools show significant promise for knowledge acquisition, particularly when huma…

classification, multivariate

Physical Systems

Ionosphere

This radar data was collected by a system in Goose Bay, Labrador. This system consists of a phased array of 16 high-frequency antennas with a total trans…

classification, multivariate

Physical Systems

Glass Identification

Vina conducted a comparison test of her rule-based system, BEAGLE, the nearest-neighbor algorithm, and discriminant analysis. BEAGLE is a product availab…

classification, multivariate

Physical Systems

3D Road Network (North Jutlan…

This dataset was constructed by adding elevation information to a 2D road network in North Jutland, Denmark (covering a region of 185 x 135 km^2). Elevati…

clustering, regression, sequential, text

Computer Science

Meta-data

This DataSet is about the results of Statlog project. The project performed a comparative study between Statistical, Neural and Symbolic learning algorith…

classification, multivariate

Others

Lenses

The examples are complete and noise free. The examples highly simplified the problem. The attributes do not fully describe all the factors affecting the d…

classification, multivariate

Others

Gisette

The digits have been size-normalized and centered in a fixed-size image of dimension 28x28. The original data were modified for the purpose of the feature…

classification, multivariate

Computer Science

HTRU2

HTRU2 is a data set which describes a sample of pulsar candidates collected during the High Time Resolution Universe Survey (South) [1]. Pulsars are a r…

classification, clustering, multivariate

Physical Systems

SPECTF Heart

The dataset describes diagnosing of cardiac Single Proton Emission Computed Tomography (SPECT) images. Each of the patients is classified into two categor…

classification, multivariate

Life Sciences

Musk (Version 2)

This dataset describes a set of 102 molecules of which 39 are judged by human experts to be musks and the remaining 63 molecules are judged to be non-musk…

classification, multivariate

Physical Systems

UJI Pen Characters (Version 2)

We have created the UJIpenchars2 character database by collecting samples from 60 writers at two different sites in two phases: 1st phase, 11 writers, car…

classification, multivariate, sequential

Computer Science

AutoUniv

The user first creates a classification model and then generates classified examples from it. To create a model, the following are specified: the number o…

classification, multivariate

Others

Folio

- The leaves were placed on a white background and then photographed. - The pictures were taken in broad daylight to ensure optimum light intensity.

classification, clustering, multivariate

Others

AAAI 2014 Accepted Papers

CSV format where each row is a paper and each column an attribute.

clustering, multivariate

Computer Science

QSAR biodegradation

The QSAR biodegradation dataset was built in the Milano Chemometrics and QSAR Research Group (Universit degli Studi Milano Bicocca, Milano, Italy). The r…

classification, multivariate

Others

Shuttle Landing Control

This is a tiny database. Michie reports that Burke's group used RULEMASTER to generate comprehendable rules for determining the conditions under which an…

classification, multivariate

Physical Systems

Primary Tumor

This is one of three domains provided by the Oncology Institutenthat has repeatedly appeared in the machine learning literature. (See also breast-cancer …

classification, multivariate

Life Sciences

Gas sensors for home activity…

This dataset has recordings of a gas sensor array composed of 8 MOX gas sensors, and a temperature and humidity sensor. This sensor array was exposed to b…

classification, multivariate, time-series

Computer Science

CalIt2 Building People Counts

Observations come from 2 data streams (people flow in and out of the building), over 15 weeks, 48 time slices per day (half hour count aggregates). The…

multivariate, time-series

Others

Wholesale customers

Provide all relevant information about your data set.

classification, clustering, multivariate

Business

Image Segmentation

The instances were drawn randomly from a database of 7 outdoor images. The images were handsegmented to create a classification for every pixel. Ea…

classification, multivariate

Others

Turkiye Student Evaluation

N/A

classification, clustering, multivariate

Others

BLOGGER

In this paper, we look for to recognize the causes of users tend to cyber space in Kohkiloye and Boyer Ahmad Province in Iran. Collecting information to f…

classification, multivariate

Computer Science

Arcene

ARCENE was obtained by merging three mass-spectrometry datasets to obtain enough training and test data for a benchmark. The original features indicate th…

classification, multivariate

Life Sciences

Breast Cancer Wisconsin (Diag…

Features are computed from a digitized image of a fine needle aspirate (FNA) of a breast mass. They describe characteristics of the cell nuclei present i…

classification, multivariate

Life Sciences

OPPORTUNITY Activity Recognit…

The OPPORTUNITY Dataset for Human Activity Recognition from Wearable, Object, and Ambient Sensors is a dataset devised to benchmark human activity recogni…

classification, multivariate, time-series

Computer Science

Mice Protein Expression

The data set consists of the expression levels of 77 proteins/protein modifications that produced detectable signals in the nuclear fraction of cortex. Th…

classification, clustering, multivariate

Life Sciences

EMG dataset in Lower Limb

2. Information database: 2.1. Protocol: 22 male subjects , 11 with different knee abnormalities previously diagnosed by a professional. They undergo thre…

multivariate, time-series

Computer Science

Blood Transfusion Service Cen…

To demonstrate the RFMTC marketing model (a modified version of RFM), this study adopted the donor database of Blood Transfusion Service Center in Hsin-Ch…

classification, multivariate

Business

MPI-I VISPR (Visual Privacy)

We present a dataset to address the problem of visual privacy - where users unintentionally leak private information when sharing personal images online, …

classification, flickr, multilabel, privacy, regression, scene

Vision

Bank Marketing

The data is related with direct marketing campaigns of a Portuguese banking institution. The marketing campaigns were based on phone calls. Often, more th…

classification, multivariate

Business

Online Handwritten Assamese C…

A dataset of online handwritten assamese characters by collecting samples from 45 writers is created. Each writer contributed 52 basic characters, 10 nume…

classification, multivariate, sequential

Computer Science

SMS Spam Collection

This corpus has been collected from free or free for research sources at the Internet: -> A collection of 425 SMS spam messages was manually extracted fr…

classification, clustering, domain-theory, multivariate, text

Computer Science

US Census Data (1990)

The data was collected as part of the 1990 census. There are 68 categorical attributes. This data set was derived from the USCensus1990raw data set. The…

clustering, multivariate

Social Sciences

FMA: A Dataset For Music Anal…

* Audio track (encoded as mp3) of each of the 106,574 tracks. It is on average 10 millions samples per track.* Nine audio features (consisting of 518 attr…

classification, clustering, multivariate, time-series

Computer Science

Gesture Phase Segmentation

The dataset is composed by features extracted from 7 videos with people gesticulating, aiming at studying Gesture Phase Segmentation. Each video is repres…

classification, clustering, multivariate, sequential, time-series

Others

ISOLET

This data set was generated as follows. 150 subjects spoke the name of each letter of the alphabet twice. Hence, we have 52 training examples from each sp…

classification, multivariate

Computer Science

Census Income

Extraction was done by Barry Becker from the 1994 Census database. A set of reasonably clean records was extracted using the following conditions: ((AAGE…

classification, multivariate

Social Sciences

Trains

Notes: - Additional "background" knowledge is supplied that provides a partial ordering on some of the attribute values. - We are providing this dataset…

classification, multivariate

Others

Iris

This is perhaps the best known database to be found in the pattern recognition literature. Fisher's paper is a classic in the field and is referenced fre…

classification, multivariate

Life Sciences

Steel Plates Faults

Type of dependent variables (7 Types of Steel Plates Faults): 1.Pastry 2.Z_Scratch 3.K_Scatch 4.Stains 5.Dirtiness 6.Bumps 7.Other_Faults

classification, multivariate

Physical Systems

ICU

Please see documentation

multivariate, time-series

Life Sciences

Heterogeneity Activity Recogn…

The Heterogeneity Dataset for Human Activity Recognition from Smartphone and Smartwatch sensors consists of two datasets devised to investigate sensor het…

classification, clustering, multivariate, time-series

Computer Science

Horse Colic

2 data files: -- horse-colic.data: 300 training instances -- horse-colic.test: 68 test instances Possible class attributes: 24 (whether lesi…

classification, multivariate

Life Sciences

Syskill and Webert Web Page R…

The HTML source of a web page is given. Users looked at each web page and inidated on a 3 point scale (hot medium cold) 50-100 pages per domain. However, …

classification, multivariate, text

Computer Science

Wilt

This data set contains some training and testing data from a remote sensing study by Johnson et al. (2013) that involved detecting diseased trees in Quick…

classification, multivariate

Life Sciences

Cervical cancer (Risk Factors)

The dataset was collected at 'Hospital Universitario de Caracas' in Caracas, Venezuela. The dataset comprises demographic information, habits, and histori…

classification, multivariate

Life Sciences

Dorothea

Drugs are typically small organic molecules that achieve their desired activity by binding to a target site on a receptor. The first step in the discovery…

classification, multivariate

Life Sciences

Ozone Level Detection

For a list of attributes, please refer to those two .names files. They use the following naming convention: All the attribute start with T means the tem…

classification, multivariate, sequential, time-series

Physical Systems

Daphnet Freezing of Gait

The Daphnet Freezing of Gait Dataset is a dataset devised to benchmark automatic methods to recognize gait freeze from wearable acceleration sensors plac…

classification, multivariate, time-series

Life Sciences

Tic-Tac-Toe Endgame

This database encodes the complete set of possible board configurations at the end of tic-tac-toe games, where "x" is assumed to have played first. The t…

classification, multivariate

Games

Dermatology

This database contains 34 attributes, 33 of which are linear valued and one of them is nominal. The differential diagnosis of erythemato-squamous diseas…

classification, multivariate

Life Sciences

gene expression cancer RNA-Seq

Samples (instances) are stored row-wise. Variables (attributes) of each sample are RNA-Seq gene expression levels measured by illumina HiSeq platform.

classification, clustering, multivariate

Life Sciences

Chess (King-Rook vs. King-Paw…

The dataset format is described below. Note: the format of this database was modified on 2/26/90 to conform with the format of all the other databases in…

classification, multivariate

Games

Activity Recognition system b…

This dataset represents a real-life benchmark in the area of Activity Recognition applications, as described in [1]. The classification tasks consist in…

classification, multivariate, sequential, time-series

Computer Science

Dodgers Loop Sensor

This loop sensor data was collected for the Glendale on ramp for the 101 North freeway in Los Angeles. It is close enough to the stadium to see unusual t…

multivariate, time-series

Others

IPUMS Census Database

The original source for this data set is the IPUMS project (RugglesSobek, 1997). The IPUMS project is a large collection of federal census data which has …

multivariate

Social Sciences

Breast Cancer Wisconsin (Orig…

Samples arrive periodically as Dr. Wolberg reports his clinical cases. The database therefore reflects this chronological grouping of the data. This group…

classification, multivariate

Life Sciences

Liver Disorders

The first 5 variables are all blood tests which are thought to be sensitive to liver disorders that might arise from excessive alcohol consumption. Each l…

multivariate

Life Sciences

Amazon Access Samples

This is a sparse data set, less than 10% of the attributes are used for each sample. The link is to a '*.tgz' file which contains two files: [amzn-anon-ac…

causal-discovery, clustering, domain-theory, regression, time-series

Business

Yeast

Predicted Attribute: Localization site of protein. ( non-numeric ). The references below describe a predecessor to this dataset and its development. They…

classification, multivariate

Life Sciences

Abalone

Predicting the age of abalone from physical measurements. The age of abalone is determined by cutting the shell through the cone, staining it, and counti…

classification, multivariate

Life Sciences

Firm-Teacher_Clave-Direction_…

The data consist of 16 binary inputs and one 'four-bit' one-hot classification output. The 16-bit inputs are binary-valued attack-point vectors. 1 indicat…

classification, multivariate

Others

Pioneer-1 Mobile Robot Data

The data were collected over a series of specifically designed trials. Our hope was to cover most of the types of sensory interactions that a Pioneer migh…

multivariate, time-series

Computer Science

HIV-1 protease cleavage

Past Usage: (a) Rgnvaldsson, You and Garwicz (2015) 'State of the art prediction of HIV-1 protease cleavage sites', Bioinformatics, vol 31 (8), p…

classification, multivariate

Life Sciences

Weight Lifting Exercises moni…

Velloso, E.; Bulling, A.; Gellersen, H.; Ugulino, W.; Fuks, H. Qualitative Activity Recognition of Weight Lifting Exercises. Proceedings of 4th Internatio…

classification, multivariate

Physical Systems

Zoo

A simple database containing 17 Boolean-valued attributes. The "type" attribute appears to be the class attribute. Here is a breakdown of which animals …

classification, multivariate

Life Sciences

Nursery

Nursery Database was derived from a hierarchical decision model originally developed to rank applications for nursery schools. It was used during several …

classification, multivariate

Social Sciences

Breast Cancer

This is one of three domains provided by the Oncology Institute that has repeatedly appeared in the machine learning literature. (See also lymphography an…

classification, multivariate

Life Sciences

Activities of Daily Living (A…

This dataset comprises information regarding the ADLs performed by two users on a daily basis in their own homes. This dataset is composed by two instanc…

classification, clustering, multivariate, sequential, time-series

Computer Science

Gas sensor arrays in open sam…

Number of instances: 18000 times-series measurements recorded from a 72 metal-oxide gas sensor array-based chemical detection platform. Number of attribu…

classification, multivariate, time-series

Computer Science

PEMS-SF

We have downloaded 15 months worth of daily data from the California Department of Transportation PEMS website, [Web Link], The data describes the occupan…

classification, multivariate, time-series

Computer Science

Relative location of CT slice…

The data was retrieved from a set of 53500 CT images from 74 different patients (43 male, 31 female). Each CT slice is described by two histograms in pol…

domain-theory, regression

Computer Science

SPECT Heart

The dataset describes diagnosing of cardiac Single Proton Emission Computed Tomography (SPECT) images. Each of the patients is classified into two categor…

classification, multivariate

Life Sciences

Audiology (Original)

This database does NOT use a standard set of attributes per instance. Contact Ray Bareiss (rbareiss '@' uunet.uucp ?) for more information. Domain exper…

classification, multivariate

Life Sciences

Quadruped Mammals

The file animals.c is a data generator of structured instances representing quadruped animals as used by Gennari, Langley, and Fisher (1989) to evaluate t…

classification, generator, multivariate

Life Sciences

Detect Malacious Executable(A…

TRAINING File : I have created training file with 100+ non malacious examples and 250+ malacious samples. NON-MALACIOUS dataset is represented by +1 while…

classification, multivariate

Computer Science

Acute Inflammations

The main idea of this data set is to prepare the algorithm of the expert system, which will perform the presumptive diagnosis of two diseases of urinary …

classification, multivariate

Life Sciences

Balloons

There are four data sets representing different conditions of an experiment. All have the same attributes. a. adult-stretch.data Inflated is true if age…

classification, multivariate

Social Sciences

Reuters RCV1 RCV2 Multilingua…

Uncompressing rcv1rcv2aminigoutte.tar.bz2 will create a directory that contains 5 subdirectories EN, FR, GR, IT and SP, corresponding to the 5 languages.…

classification, multivariate

Life Sciences

Optical Recognition of Handwr…

We used preprocessing programs made available by NIST to extract normalized bitmaps of handwritten digits from a preprinted form. From a total of 43 peopl…

classification, multivariate

Computer Science

HEPMASS

Machine learning is used in high-energy physics experiments to search for the signatures of exotic particles. These signatures are learned from Monte Carl…

classification, multivariate

Physical Systems

Sales_Transactions_Dataset_We…

52 columns for 52 weeks; normalised values of provided too.

clustering, multivariate, time-series

Others

Thyroid Disease

# From Garavan Institute # Documentation: as given by Ross Quinlan # 6 databases from the Garavan Institute in Sydney, Australia # Approximately the follo…

classification, domain-theory, multivariate

Life Sciences

banknote authentication

Data were extracted from images that were taken from genuine and forged banknote-like specimens. For digitization, an industrial camera usually used for …

classification, multivariate

Computer Science

Adult

Extraction was done by Barry Becker from the 1994 Census database. A set of reasonably clean records was extracted using the following conditions: ((AAGE…

classification, multivariate

Social Sciences

Teaching Assistant Evaluation

The data consist of evaluations of teaching performance over three regular semesters and two summer semesters of 151 teaching assistant (TA) assignments a…

classification, multivariate

Others

Waveform Database Generator (…

Notes: -- 3 classes of waves -- 21 attributes, all of which include noise -- See the book for details (49-55, 169) -- waveform.data.Z …

classification, generator, multivariate

Physical Systems

Website Phishing

The phishing problem is considered a vital issue in .COM industry especially e-banking and e-commerce taking the number of online transactions involving p…

classification, multivariate

Computer Science

Internet Advertisements

This dataset represents a set of possible advertisements on Internet pages. The features encode the geometry of the image (if available) as well as phras…

classification, multivariate

Computer Science

Homepage

Description

Tags

Discussion

Related datasets