Landmark 3D

Yahoo Flickr Creative Commons…

Yahoo Flickr Creative Commons 100M (YFCC100M) dataset contains a list of photos and videos. This list is compiled from data available on Yahoo! Flickr. Al…

3d, clustering, community, detection, flickr, image, internet, landmark, recognition, reconstruction, social

Vision

ETHZ CVL RueMonge 2014

This ETHZ CVL RueMonge 2014 dataset used for 3D reconstruction and semantic mesh labelling for urban scene understanding. It was first published in [1] …

3d, architecture, benchmark, classification, code, mesh, outdoor, paris, pointcloud, recognition, reconstruction, segmentation, semantic, source, urban

Vision

ISPRS Urban Classification

ISPRS Test Project on Urban Classification and 3D Building Reconstruction The ISPRS working group III/4 announces the release of the 2D semantic labelin…

3d, building, city, classification, recognition, reconstruction, semantic, urban

Vision

Yotta

The Yotta dataset consists of 70 images for semantic labeling given in 11 classes. It also contains multiple videos and camera matrices for 14km or drivin…

3d, camera, classification, reconstruction, segmentation, semantic, urban, video

Vision

B3DO: Berkeley 3D Object Data…

For the first few decades of the fields existence, computer vision has been focused on algorithmic, logical approaches to perception. But it was only with…

3d, depth, indoor, kinect, object, recognition, reconstruction

Vision

Pittsburgh Fast-food Image da…

The Pittsburgh Fast-food Image dataset (PFID) consists of 4545 still images, 606 stereo pairs, 3033600 videos for structure from motion, and 27 privacy-pr…

classification, food, laboratory, real, recognition, reconstruction, video

Vision

Visual Search Patches

The Compact Descriptors for Visual Search Patches Dataset (CDVS) is a dataset comprised of pairwise image patches. MPEG is a standard titled Compact Desc…

descriptor, feature, matching, mpeg, patch, retrieval

Vision

Labeling in 3D Scenes

This dataset package contains the software and data used for Detection-based Object Labeling on the RGB-D Scenes Dataset as implemented in the paper: De…

3d, depth, indoor, kinect, object, recognition, reconstruction

Vision

SHOT 3D shape description

The 3D shape description dataset consists of multiple sub-datasets Descriptor Matching - Dataset 1 & 2 (Stanford) These datasets, created from some of …

3d, benchmark, description, matching, reconstruction, registration, shape

Vision

Landmark 1000

The Landmark 1000 or 1k dataset is a collection of the top 1000 popular flickr landmarks mined from flickr. It is maintained by Noah Snavely and publish…

3d, estimation, landmark, location, pointcloud, pose, reconstruction, world

Vision

3DVis

The 3DVis dataset includes a set of 12 heterogeneous scenes for testing 3D scene registration and analysis methods. Models include homogeneous shapes, rep…

3d, matching, reconstruction, registration, shape, symmetry

Vision

1DSfM Landmarks

The 1DSfM Landmarks is a collection of community-based image reconstruction by Kyle Wilson and is comprised of 14 datasets with comparison to bundler grou…

3d, benchmark, city, groundtruth, landmark, reconstruction, urban

Vision

Notre Dame

The Notre Dame de Paris dataset used for 3D SfM reconstruction and contains 715 images provided by Noah Snavely. There are also version for NotreDame b…

3d, 3d reconstruction, flickr, frontview, landmark, limited, paris, pointcloud, sfm

Vision

MSR Action

The MSR Action datasets is a collection of various 3D datasets for action recognition. See details http://research.microsoft.com/en-us/um/people/zliu/a…

3d, action, detection, recognition, reconstruction, video

Vision

CMP WxBS dataset

The Wide (multiple) Baseline Dataset. 31 image pairs, simultaneously combining several nuisance factors: geometry, illumination, IR-visible, etc. WxBS: …

day, description, detection, feature, ir, matching, night, viewpoint

Vision

ISPRS-EuroSDR HighDensity

ISPRS and EuroSDR - Benchmark on High Density Aerial Image Matching Background and Scope of the project Innovations in matching algorithms as well as …

3d, aerial, benchmark, city, germany, multiview, photogrammetry, reconstruction, switzerland, urban

Vision

Robotic 3D Scan Repository

The Robotic 3D Scan Repository from Osnabrueck contains 23 different datasets showing a veriaty of 3D scans for objects, humans, cities, university campus…

3d, aerial, bremen, city, germany, heat, human, laser, lidar, osnabrueck, reconstruction, scan, urban

Vision

Leuven Stereo Scene

The Leuven Stereo Scene dataset is a scene and depth dataset. There exist two variants of this dataset - a CVPR 2007 paper [1] by Leibe et al. for detecti…

3d, depth, leuven, reconstruction, segmentation, semantic, sfm, stereo, urban

Vision

Fish4Knowledge

The Fish4Knowledge project (groups.inf.ed.ac.uk/f4k/) is pleased to announce the availability of 2 subsets of our tropical coral reef fish video and ext…

animal, camera, classification, fish, motion, nature, recognition, video, water

Vision

Visual Attributes dataset

The Visual Attributes dataset contains visual attribute annotations for over 500 object classes (animate and inanimate) which are all represented in Image…

attribute, classification, imagenet, object, recognition

Vision

AWS Public Datasets

AWS hosts a variety of public datasets that anyone can access for free. Previously, large datasets such as satellite imagery or genomic data have require…

amazon, biology, classification, deep, human, image, learning, recognition, resolution, satellite, segmentation, space

Vision

SceneNet RGB-D Synthetic Indo…

SceneNet RGB-D is dataset comprised of 5 million Photorealistic Images of Synthetic Indoor Trajectories with Ground Truth. It expands the previous work of…

3d, indoor, lighting, navigation, reconstruction, rendering, robot, scene, segmentation, slam, synthetic, trajectory

Vision

Farman Institute 3D Point Sets

The Farman Institute 3D Point Sets dataset contains 11 objects by a 3D laser scanner. This dataset was peer-reviewed by Image Processing On Line: Farman I…

3d, laser, model, object, point, reconstruction, scanner

Vision

Paris500k

The Paris500k dataset consists of 501,356 geotagged images collected from Flickr and Panoramio. The dataset was collected from a geographic bounding box r…

3d reconstruction, flickr, geotag, image retrieval, landmark, panoramio, paris, sfm

Vision

CMP Extreme View Dataset

15 wide baseline stereo image pairs with large viewpoint change, provided ground truth homographies. Image size (~1000x700 pixels, RGB) D. Mishkin and…

description, detection, feature, matching, viewpoint, wide baseline stereo

Vision

The Oxford RobotCar Dataset

The Oxford RobotCar Dataset contains over 100 repetitions of a consistent route through Oxford, UK, captured over a period of over a year. The dataset cap…

autonomous, car, classification, detection, driving, recognition, robot, segmentation, street, time, urban, video, year

Vision

Colosseum and San Marco

The Colosseum and San Marco are two image datasets for dense multiview stereo reconstructions used for evaluating the visual photo realism. The datasets…

3d reconstruction, aerial, flickr, landmark, photo-realism, sfm, streetside, urban

Vision

50 Salads

The dataset captures 25 people preparing 2 mixed salads each and contains over 4h of annotated accelerometer and RGB-D video data. Annotated activities co…

action, activity, classification, detection, recognition, tracking, video

Vision

New College Data

The New College Data Set contains 30GB of data intended for use by the mobile robotics and vision research communities. Our anticipated users are parties …

3d, navigation, odometry, panorama, path, reconstruction, stereo, urban

Vision

VIDEO datasets overview

Many different labeled video datasets have been collected over the past few years, but it is hard to compare them at a glance. So we have created a handy …

action, benchmark, classification, detection, object, recognition, video

Vision

CASIA Gait Recognition Dataset

Dataset A (former NLPR Gait Database) was created on Dec. 10, 2001, including 20 persons. Each person has 12 image sequences, 4 sequences for each of the …

action, biometry, classification, foot, gait, human, motion, pressure, recognition

Vision

HCI 4D Lightfields

The HCI 4D Lightfields dataset contains 11 objects with corresponding lightfields for depth estimation. Datasets can be downloaded individually below. F…

3d, 4d, benchmark, depth, evaluation, lightfield, reconstruction

Vision

Osnabruck - Synthetic Scalabl…

Voxel Based Dataset for Systematic 3D reconstruction by artificial neural networks (ANNs). A synthetic scalable cube dataset for training, testing and v…

3d, deep learning, reconstruction, sfm, synthetic city urban

Vision

SUNCG: Indoor Scenes

The SUNCG dataset is a Large 3D Model Repository for Indoor Scenes. SUNCG is an ongoing effort to establish a richly-annotated, large-scale dataset of…

3d, indoor, layout, object, realism, recognition, rendering, room, scene, segmentation, synthetic

Vision

San Francisco Landmark Datase…

The San Francisco Landmark Dataset for Mobile Landmark Recognition is a set of images and query images for localization. We present the San Francisco La…

calibration, city, gps, landmark, localization, mobile, retrieval, sanfrancisco, urban

Vision

RGB-D Person Re-identification

The RGB-D Person Re-identification dataset is for person re-identification using depth information. The main motivation is that the standard techniques (s…

3d, classification, depth, identification, pedestrian, shape

Vision

Berkeley Multimodal Human Act…

The Berkeley Multimodal Human Action Database (MHAD) contains 11 actions performed by 7 male and 5 female subjects in the range 23-30 years of age except …

action, classification, motion, multiview, recognition

Vision

MPI-I VISPR (Visual Privacy)

We present a dataset to address the problem of visual privacy - where users unintentionally leak private information when sharing personal images online, …

classification, flickr, multilabel, privacy, regression, scene

Vision

Matterport 2D-3D-Semantics Da…

The 2D-3D-S dataset provides a variety of mutually registered modalities from 2D, 2.5D and 3D domains, with instance-level semantic and geometric annotati…

3d, building, depth, indoor, large-scale, normal, panorama, reconstruction, segmentation, semantic

Vision

MIT Places205

Places205 dataase contains 2.5 million images from 205 scene categories for the academic public. The image dataset contains 2,448,873 images from 205 sc…

feature, learning, place, recognition, scene, urban

Vision

INRIA Lafarge Benchmarks

Some datasets and evaluation tools are provided on this page for four different computer vision and computer graphics problems. Population counting Lin…

3d, counting, crowd, detection, groundtruth, line, network, object, pedestrian, pointcloud, reconstruction, road, surface, urban

Vision

Caltech Game Covers Dataset

The Caltech Game Covers dataset consists of CD/DVD covers of video games. The set was downloaded from freecovers.net during the summer of 2008. The set in…

caltech, classification, cover, game, hierarchy, retrieval, taxonomy

Vision

ISPRS WG III/4

ISPRS Test Project on Urban Classification, 3D Building Reconstruction and Semantic Labeling. In this part of our working group site you will get further …

3d, aerial, benchmark, canada, city, germany, multiview, photogrammetry, recognition, segmentation, semantic, urban

Vision

CMP map2photo

The CMP map2photo dataset consists of 6 pairs, where one image is satellite photo and second image is a map of the same area. The task is to match these …

baseline, description, detection, feature, map, matching, remote, sensing, wide

Vision

ISPRS-EuroSDR Multi-Platform

ISPRS / EuroSDR Benchmark for Multi-Platform Photogrammetry In these pages you can get information about the BENCHMARK FOR MULTI-PLATFORM PHOTOGRAMMETRY…

3d, aerial, benchmark, city, germany, multiview, photogrammetry, reconstruction, switzerland, urban

Vision

Paris Rue Madame

Paris-rue-Madame dataset contains 3D Mobile Laser Scanning (MLS) data from rue Madame, a street in the 6th Parisian district (France). The test zone conta…

3d, classification, laser, pointcloud, segmentation, semantic

Vision

Symmetry Set

The Symmetry set dataset is a collection of images at different illuminations for the purpose of image matching using local symmetry features. Image Mat…

building, feature, illumination, image, lighting, matching, symmetry, urban

Vision

ScanNet

ScanNet is an RGB-D video dataset containing 2.5 million views in more than 1500 scans, annotated with 3D camera poses, surface reconstructions, and insta…

3d, cad, indoor, layout, object, realism, recognition, rendering, room, scene, segmentation, synthetic

Vision

UBO 2014 Materials

The UBO 2014 consists of 7 semantic categories. Each of these 7 material categories contains measurements of 12 different material instances for being cap…

classification, illumination, light, material, recognition, texture

Vision

Webcam Interestingness

The Webcam Interestingness dataset consists of 20 different webcam streams, with 159 images each. It is annotated with interestingness ground truth, acqui…

classification, interest, ranking, retrieval, video, weather, webcam

Vision

CMP Extreme Zoom Dataset

The Extreme Zoom Dataset. EZD is a 6 image sets with incleasing zoom factor from general scene view to focusing on single detail. MODS: Fast and Robust …

description, detection, feature, matching, viewpoint, zoom

Vision

FlickrLogos-32

The FlickrLogos-32 dataset contains photos showing brand logos and is meant for the evaluation of multi-class logo recognition as well as logo retrieval m…

classification brand boundingbox, detection, flickr, image, logo, machine learning, object recognition, retrieval

Vision

Comprehensive Cars (CompCars)

The Comprehensive Cars (CompCars) dataset contains data from two scenarios, including images from web-nature and surveillance-nature. The web-nature data …

attribute, car, classification, fine-grained, object, recognition, urban, vehicle

Vision

Stable Structure from Motion

The Stable Structure from Motion datasets due to size limitations cannot put the images online. Instead here are the tracked image points and the final re…

3d, 3d reconstruction, church, geometry, landmark, robust, sfm, stability

Vision

FGVC-Aircraft

Fine-Grained Visual Classification of Aircraft (FGVC-Aircraft) is a benchmark dataset for the fine grained visual categorization of aircraft. Data, anno…

aircraft, airplane, benchmark, classification, evaluation, fine-grained, recognition

Vision

Symmetry Facades

The Symmetry Facades dataset contains 9 building facades with multiple images. It used for coupled symmetry and structure from motion detection. Couple…

3d, building, facade, reconstruction, repetition, sfm, symmetry, urban

Vision

MPI VehicleScenes

Abstract Scene understanding has (again) become a focus of computer vision research, leveraging advances in detection, context modeling, and tracking. In…

3d, car, classification, pedestrian, scene, segmentation, semantic, understanding

Vision

3D Mask Attack Dataset

The 3D Mask Attack Database (3DMAD) is a biometric (face) spoofing database. It currently contains 76500 frames of 17 persons, recorded using Kinect for b…

3d, biometry, emotion, face, frontview, recognition, segmentation

Vision

udacity self-driving-car

At Udacity, we believe in democratizing education. How can we provide opportunity to everyone on the planet? We also believe in teaching really amazing an…

autonomous, car, classification, detection, driving, recognition, robot, segmentation, street, synthetic, time, urban, video

Vision

Street View House Number (SVH…

SVHN is a real-world image dataset for developing machine learning and object recognition algorithms with minimal requirement on data preprocessing and fo…

classification, detection, number, real, recognition, streetside, streetview, text, urban, world

Vision

UMD Dynamic Scene Recognition

The UMD Dynamic Scene Recognition dataset consists of 13 classes and 10 videos per class and is used to classify dynamic scenes. The dataset has been de…

classification, dynamic, motion, recognition, scene, video

Vision

CMP Facades

The CMP Facade dataset consists of facade images assembled at the Center for Machine Perception, which includes 600 rectified images of facades from vario…

classification, facade, recognition, rectification, segmentation, semantic, similarity, structure, urban

Vision

Molecular Biology (Protein Se…

This is a data set used by Ning Qian and Terry Sejnowski in their study using a neural net to predict the secondary structure of certain globular proteins…

classification, sequential

Life Sciences

Historical Car Database

The database contains historical car images from 1920s to 1990s crawled from cardatabase.net. There are 10130 training and 3343 test images. Annotations…

car, recognition, time

Vision

Nursery

Nursery Database was derived from a hierarchical decision model originally developed to rank applications for nursery schools. It was used during several …

classification, multivariate

Social Sciences

Breast Cancer

This is one of three domains provided by the Oncology Institute that has repeatedly appeared in the machine learning literature. (See also lymphography an…

classification, multivariate

Life Sciences

3D Object in Clutter Recognit…

The dataset is composed of 150 synthetic scenes, captured with a (perspective) virtual camera, and each scene contains 3 to 5 objects. The model set is co…

mesh, recognition, segmentation, synthetic

Vision

Activities of Daily Living (A…

This dataset comprises information regarding the ADLs performed by two users on a daily basis in their own homes. This dataset is composed by two instanc…

classification, clustering, multivariate, sequential, time-series

Computer Science

Depth-Based Person Identifica…

Depth-Based Person Identification from Top View Dataset.

recognition

Vision

Gas sensor arrays in open sam…

Number of instances: 18000 times-series measurements recorded from a 72 metal-oxide gas sensor array-based chemical detection platform. Number of attribu…

classification, multivariate, time-series

Computer Science

PEMS-SF

We have downloaded 15 months worth of daily data from the California Department of Transportation PEMS website, [Web Link], The data describes the occupan…

classification, multivariate, time-series

Computer Science

SPECT Heart

The dataset describes diagnosing of cardiac Single Proton Emission Computed Tomography (SPECT) images. Each of the patients is classified into two categor…

classification, multivariate

Life Sciences

Audiology (Original)

This database does NOT use a standard set of attributes per instance. Contact Ray Bareiss (rbareiss '@' uunet.uucp ?) for more information. Domain exper…

classification, multivariate

Life Sciences

Quadruped Mammals

The file animals.c is a data generator of structured instances representing quadruped animals as used by Gennari, Langley, and Fisher (1989) to evaluate t…

classification, generator, multivariate

Life Sciences

Breast Cancer Wisconsin (Prog…

Each record represents follow-up data for one breast cancer case. These are consecutive patients seen by Dr. Wolberg since 1984, and include only those c…

classification, multivariate, regression

Life Sciences

Ljubljana CVL Face Database

Database contains 798 images of 114 persons, with 7 images per person and is freely available for research purposes. All images were taken in supervised c…

biometry, face, human, illumination, lighting, pedestrian, person, recognition

Vision

Detect Malacious Executable(A…

TRAINING File : I have created training file with 100+ non malacious examples and 250+ malacious samples. NON-MALACIOUS dataset is represented by +1 while…

classification, multivariate

Computer Science

Parkinson Speech Dataset with…

The PD database consists of training and test files. The training data belongs to 20 PWP (6 female, 14 male) and 20 healthy individuals (10 female, 10 mal…

classification, multivariate, regression

Life Sciences

TST fall detection

It is composed of ADL (activity daily living) and fall actions simulated by 11 volunteers. The people involved in the test are aged between 22 and 39, wit…

accelerometer, action, depth, fall detection - adl, human, kinect, recognition, video, wearable

Vision

KITTI Odometry

http://www.cvlibs.net/datasets/kitti/eval_odometry.php Related Datasets TUM RGB-D Dataset: Indoor dataset captured with Microsoft Kinect and high-accu…

localization, matching, navigation, odometry, registration, slam, urban path 3d reconstruction

Vision

Acute Inflammations

The main idea of this data set is to prepare the algorithm of the expert system, which will perform the presumptive diagnosis of two diseases of urinary …

classification, multivariate

Life Sciences

Balloons

There are four data sets representing different conditions of an experiment. All have the same attributes. a. adult-stretch.data Inflated is true if age…

classification, multivariate

Social Sciences

SIID

The SIID silhouette dataset contains... and is from the Shape Indexing of Image Database (SIID). Download SIID silhouette dataset http://www.lems.brown…

binary, matching, shape retrieval

Vision

Reuters RCV1 RCV2 Multilingua…

Uncompressing rcv1rcv2aminigoutte.tar.bz2 will create a directory that contains 5 subdirectories EN, FR, GR, IT and SP, corresponding to the 5 languages.…

classification, multivariate

Life Sciences

MOT Challenge 2D and 3D

The MOT Challenge is a framework for the fair evaluation of multiple people tracking algorithms. In this framework we provide: - A large collection of d…

3d, benchmark, benhttp://motchallenge.net/chmark, dataset, evaluation, multiple, pedestrian, people, surveillance, target, tracking, video

Vision

Optical Recognition of Handwr…

We used preprocessing programs made available by NIST to extract normalized bitmaps of handwritten digits from a preprinted form. From a total of 43 peopl…

classification, multivariate

Computer Science

HEPMASS

Machine learning is used in high-energy physics experiments to search for the signatures of exotic particles. These signatures are learned from Monte Carl…

classification, multivariate

Physical Systems

Caltech 101

Pictures of objects belonging to 101 categories

classification, detection

Vision

Mythological Creatures

The Mythological Creatures consists of articulated shapes (silhouettes) for partial similarity experiments and contains 15 shapes: 5 humans, 5 horses and …

animal, binary, matching, partial, shape retrieval

Vision

GPS Trajectories

The dataset is composed by two tables. The first table go_track_tracks presents general attributes and each instance has one trajectory that is represente…

classification, multivariate, regression

Computer Science

CMU Face Images

Each image can be characterized by the pose, expression, eyes, and size. There are 32 images for each person capturing every combination of features. To…

classification, image

Others

FaceScrub

The FaceScrub dataset comprises a total of 107818 unconstrained face images of 530 celebrities crawled from the Internet, with about 200 images per person…

celebrity, detection, face, human, people, recognition

Vision

Thyroid Disease

# From Garavan Institute # Documentation: as given by Ross Quinlan # 6 databases from the Garavan Institute in Sydney, Australia # Approximately the follo…

classification, domain-theory, multivariate

Life Sciences

banknote authentication

Data were extracted from images that were taken from genuine and forged banknote-like specimens. For digitization, an industrial camera usually used for …

classification, multivariate

Computer Science

Adult

Extraction was done by Barry Becker from the 1994 Census database. A set of reasonably clean records was extracted using the following conditions: ((AAGE…

classification, multivariate

Social Sciences

Teaching Assistant Evaluation

The data consist of evaluations of teaching performance over three regular semesters and two summer semesters of 151 teaching assistant (TA) assignments a…

classification, multivariate

Others

Waveform Database Generator (…

Notes: -- 3 classes of waves -- 21 attributes, all of which include noise -- See the book for details (49-55, 169) -- waveform.data.Z …

classification, generator, multivariate

Physical Systems

Cargo 2000 Freight Tracking a…

A description of the underlying Cargo 2000 standard and the processes reflected in the data set can be found at [Web Link].

classification, multivariate, regression, sequential

Business

Website Phishing

The phishing problem is considered a vital issue in .COM industry especially e-banking and e-commerce taking the number of online transactions involving p…

classification, multivariate

Computer Science

Internet Advertisements

This dataset represents a set of possible advertisements on Internet pages. The features encode the geometry of the image (if available) as well as phras…

classification, multivariate

Computer Science

UJIIndoorLoc

Many real world applications need to know the localization of a user in the world to provide their services. Therefore, automatic user localization has be…

classification, multivariate, regression

Computer Science

Berkeley Urban Street tracking

The UrbanStreet dataset used in the paper can be downloaded here [188M] . It contains 18 stereo sequences of pedestrians taken from a stereo rig mounted o…

detection, human, multitarget, pedestrian, recognition, segmentation, tracking, urban, video

Vision

Robot Execution Failures

The donation includes 5 datasets, each of them defining a different learning problem: * LP1: failures in approach to grasp position * LP2: failur…

classification, multivariate, time-series

Physical Systems

Anuran Calls (MFCCs)

This dataset was used in several classifications tasks related to the challenge of anuran species recognition through their calls. It is a multilabel data…

classification, clustering, multivariate

Life Sciences

seismic-bumps

Mining activity was and is always connected with the occurrence of dangers which are commonly called mining hazards. A special case of such threat is a s…

classification, multivariate

Others

Polish companies bankruptcy d…

The dataset is about bankruptcy prediction of Polish companies. The data was collected from Emerging Markets Information Service (EMIS, [Web Link]), which…

classification, multivariate

Business

Sentiment Labelled Sentences

This dataset was created for the Paper 'From Group to Individual Labels using Deep Features', Kotzias et. al,. KDD 2015 Please cite the paper if you want …

classification, text

Others

USPTO Algorithm Challenge, ru…

USPTO Algorithm Challenge, run by NASA-Harvard Tournament Lab and TopCoder Problem: Patent Labeling

classification, domain-theory

Others

CMLA Subpixel Stereo Dataset

A 66 stereo pairs dataset with their subpixel ground truths. The construction and improvement of algorithms for subpixel stereovision requires very prec…

3d, depth, groundtruth, noise, pointcloud, stereo, stereovision, subpixel

Vision

SECOM

A complex modern semi-conductor manufacturing process is normally under consistent surveillance via the monitoring of signals/variables collected from sen…

causal-discovery, classification, multivariate

Computer Science

Phishing Websites

One of the challenges faced by our research was the unavailability of reliable training datasets. In fact this challenge faces any researcher in the field…

classification

Computer Security

Caltech

Cars, Motorcycles, Airplanes, Faces, Leaves, Backgrounds

classification, detection

Vision

KDD Cup 1999 Data

Please see task description.

classification, multivariate

Computer Science

DTU Robot

The DTU Robot dataset consists of color images of 60 scenes acquired in a controlled setup from 119 different positions and under different lighting. For …

feature description, feature detection, feature matching, illumination, reconstruction, sfm

Vision

Twin gas sensor arrays

This dataset includes the recordings of five replicates of an 8-sensor array. Each unit holds 8 MOX sensors and integrates custom-designed electronics for…

classification, domain-theory, multivariate, regression, time-series

Computer Science

Gas Sensor Array Drift Datase…

This data set contains 13,910 measurements from 16 chemical sensors exposed to 6 gases at different concentration levels. This dataset is an extension of …

causa, classification, clustering, multivariate, regression, time-series

Computer Science

ImageNET

The ImageNET dataset is the latest dataset by Li Fei-Fei containing various dataset ranging from 1000 to 10000 categories.

image classification, object segmentation, retrieval

Vision

Synthetic CAD models

The Synthetic CAD Models dataset consists of X synthetic CAD models for detection (planar) primitives. Efficient RANSAC for Point-Cloud Shape Detection …

3d object, model fitting, primitive, ransac, reconstruction, synthetic

Vision

Heart Disease

This database contains 76 attributes, but all published experiments refer to using a subset of 14 of them. In particular, the Cleveland database is the o…

classification, multivariate

Life Sciences

Wine

These data are the results of a chemical analysis of wines grown in the same region in Italy but derived from three different cultivars. The analysis dete…

classification, multivariate

Physical Systems

ZuBuD+

ZuBuD+, created in February 2017 by Federico Magliani (University of Parma), introduces many query images balancing the class evaluated from the previous …

building, image retrieval, landmark, urban

Vision

Caltech 256

Pictures of objects belonging to 256 categoriesPictures of objects belonging to 256 categories.

classification, natural-image

Vision

WWW Crowd

The Where Who Why (WWW) dataset provides 10,000 videos with over 8 million frames from 8,257 diverse scenes, therefore offering a superior comprehensive d…

crowd, detection, flow, optical, pedestrian, recognition, surveillance, video

Vision

Quad 6K

The Quad 6K dataset is a Structure-from-Motion dataset taken at Arts Quad at Cornell University campus and consists of 6514 images with ground truth posit…

3d gps, 3d reconstruction, groundtruth, landmark, sfm, urban

Vision

GaTech VideoContext

The GaTech VideoContext dataset consists of over 100 groundtruth annotated outdoor videos with over 20000 frames for the task of geometric context evalua…

classification, context, geometry, nature, outdoor, segmentation, semantic, supervised, unsupervised, urban, video

Vision

User Knowledge Modeling

-- The users' knowledge class were classified by the authors using intuitive knowledge classifier (a hybrid ML technique of k-NN and meta-heuristic…

classification, clustering, multivariate

Computer Science

Ford Car Dataset

The Ford Car dataset is joint effort of Pandey et al. (for collecting images, Lidar points, calibration etc.) and us (for annotation of 2D and 3D objects)…

3d, car, detection, groundtruth, lidar, sfm

Vision

Echocardiogram

All the patients suffered heart attacks at some point in the past. Some are still alive and some are not. The survival and still-alive variables, when ta…

classification, multivariate

Life Sciences

Letter Recognition

The objective is to identify each of a large number of black-and-white rectangular pixel displays as one of the 26 capital letters in the English alphabet…

classification, multivariate

Computer Science

Madelon

MADELON is an artificial dataset containing data points grouped in 32 clusters placed on the vertices of a five dimensional hypercube and randomly labeled…

classification, multivariate

Others

Haberman's Survival

The dataset contains cases from a study that was conducted between 1958 and 1970 at the University of Chicago's Billings Hospital on the survival of patie…

classification, multivariate

Life Sciences

ETHZ Shape

The ETHZ Shape classes dataset from Vittorio Ferrari [?] consists of five object classes and a total of 255 images. All classes contain significant intra-…

applelogo, bottle, clutter, giraffe, matching, mug, nature, object detection by shape, segmentation, swan

Vision

Breast Tissue

Impedance measurements were made at the frequencies: 15.625, 31.25, 62.5, 125, 250, 500, 1000 KHz Impedance measurements of freshly excised breast tissue …

classification, multivariate

Life Sciences

Wine Quality

The two datasets are related to red and white variants of the Portuguese "Vinho Verde" wine. For more details, consult: [Web Link] or the reference [Corte…

classification, multivariate, regression

Business

Waveform Database Generator (…

Notes: -- 3 classes of waves -- 40 attributes, all of which include noise -- The latter 19 attributes are all noise attributes with mean…

classification, generator, multivariate

Physical Systems

Student Performance

This data approach student achievement in secondary education of two Portuguese schools. The data attributes include student grades, demographic, social a…

classification, multivariate, regression

Social Sciences

Volcanoes on Venus - JARtool …

The data was collected by the Magellan spacecraft over an approximately four year period from 1990--1994. The objective of the mission was to obtain globa…

classification, image

Physical Systems

Microsoft COCO

The Microsoft COCO (mscoco) is an image recognition and segmentation dataset which contains more 300k images for more than 70 categories. Other features…

benchmark, context, detection, object, recognition, segmentation, semantic

Vision

Statlog (Australian Credit Ap…

This file concerns credit card applications. All attribute names and values have been changed to meaningless symbols to protect confidentiality of the da…

classification, multivariate

Financial

Grammatical Facial Expressions

The automated analysis of facial expressions has been widely used in different research areas, such as biometrics or emotional analysis. Special import…

classification, clustering, multivariate, sequential

Computer Science

YouTube Spam Collection

The table below lists the datasets, the YouTube video ID, the amount of samples in each class and the total number of samples per dataset. Dataset --- Yo…

classification, text

Computer Science

Burst Header Packet (BHP) flo…

For Further information about the variables see the file in the data folder.

classification, text

Computer Science

Daily and Sports Activities

Brief Description of the Dataset: --------------------------------- Each of the 19 activities is performed by eight subjects (4 female, 4 male, between th…

classification, clustering, multivariate, time-series

Computer Science

Gastrointestinal Lesions in R…

This dataset contains the features extracted from a database of colonoscopic videos showing gastrointestinal lesions. It also contains the ground truth co…

classification, multivariate

Computer Science

seeds

The examined group comprised kernels belonging to three different varieties of wheat: Kama, Rosa and Canadian, 70 elements each, randomly selected for the…

classification, clustering, multivariate

Life Sciences

Multiple Instance Learning da…

MIL data sets used in our 2002 NIPS paper for Elepphant, Musk, TREC http://www.cs.cmu.edu/~juny/MILL/MIL-experiments.htm

classification, machine learning

Vision

Domain-specific Personal Vide…

The domain-specific personal videos highlight dataset from the paper [1] describes a fully automatic method to train domain-specific highlight ranker for…

action, domain, human, recognition, saliency, summarization, video, wearable

Vision

Post-Operative Patient

The classification task of this database is to determine where patients in a postoperative recovery area should be sent to next. Because hypothermia is a…

classification, multivariate

Life Sciences

Tools2D

The Tools 2D dataset from Bronstein, Bronstein, Bruckstein, and Kimmel [?] for partial similarity experiments and consists of 15 shapes: 5 humans, 5 horse…

binary, matching, partial, shape retrieval

Vision

Pen-Based Recognition of Hand…

We create a digit database by collecting 250 samples from 44 writers. The samples written by 30 writers are used for training, cross-validation and writer…

classification, multivariate

Computer Science

Annotated Web Ears Dataset (A…

Dataset contains 1000 images of 100 persons, with 10 images per person and is freely available. All images were acquired by cropping ears from images from…

biometry, ear, human, lighting, pedestrian, person, recognition

Vision

ICDAR 2011

This challenge is set up around three tasks: Text Localisation, Text Segmentation and Word Recognition. Participation in any or all tasks is welcome. Chec…

classification, text detection, text recognition

Vision

Record Linkage Comparison Pat…

The records represent individual data including first and family name, sex, date of birth and postal code, which were collected through iterative insertio…

classification, multivariate

Others

Twitter Data set for Arabic S…

--- By using a tweet crawler, we collect 2000 labelled tweets (1000 positive tweets and 1000 negative ones) on various topics such as: politics an…

classification, text

Social Sciences

NoisyOffice

AIMS AND PURPOSES This corpus is intended to do cleaning (or binarization) and enhancement of noisy grayscale printed text images using supervised learni…

classification, multivariate, regression

Computer Science

Chess (King-Rook vs. King)

An Inductive Logic Programming (ILP) or relational learning framework is assumed (Muggleton, 1992). The learning system is provided with examples of chess…

classification, multivariate

Games

MoCap Hand Postures

A Vicon motion capture camera system was used to record 12 users performing 5 hand postures with markers attached to a left-handed glove. A rigid patter…

classification, clustering, multivariate

Computer Science

Wearable Computing: Classific…

IMPORTANT: we have lower performance on 'leave-one-subject-out' tests. The performance baseline index we established is for 10-fold cross-validation tests…

classification, sequential

Computer Science

An RGB-D Dataset for 6D Pose …

A dataset acquired with 3 synchronized sensors (Primesense Carmine 1.09, Microsoft Kinect v2, Canon IXUS 950 IS), featuring: * 30 industry-relevant obje…

3d, estimation, object, pose, rgbd, texture-less

Vision

ETH CVL IMDB WIKI Faces

Since the publicly available face image datasets are often of small to medium size, rarely exceeding tens of thousands of images, and often without age in…

age, biometry, detection, face, imdb, recognition, wikipedia

Vision

MPI Multi-View Collection GVV…

Welcome to the homepage of the gvvperfcapeva datasets. This site serves as a hub to access a wide range of datasets that have been created for projects of…

action, depth, face, human, mesh, multiview, pose, reconstruction, tracking, video

Vision

Reuters Transcribed Subset

Data Characteristics: -------------------- This data was created by selecting 20 files each from the 10 largest classes in the Reuters-21578 collection …

classification, text

Business

Multiple Features

This dataset consists of features of handwritten numerals (`0'--`9') extracted from a collection of Dutch utility maps. 200 patterns per class (for a tota…

classification, multivariate

Computer Science

Soybean (Large)

There are 19 classes, only the first 15 of which have been used in prior work. The folklore seems to be that the last four classes are unjustified by the …

classification, multivariate

Life Sciences

Census-Income (KDD)

This data set contains weighted census data extracted from the 1994 and 1995 Current Population Surveys conducted by the U.S. Census Bureau. The data cont…

classification, multivariate

Social Sciences

Devanagari Handwritten Charac…

Data Type: GrayScale Image The image dataset can be used to benchmark classification algorithm for OCR systems. The highest accuracy obtained in the Test…

classification

Computer Science

NYU Depth v1

The NYU-Depth data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft Ki…

depth, kinect, label, reconstruction, semantic segmentation

Vision

Arrhythmia

This database contains 279 attributes, 206 of which are linear valued and the rest are nominal. Concerning the study of H. Altay Guvenir: "The aim is to…

classification, multivariate

Life Sciences

LED Display Domain

This simple domain contains 7 Boolean attributes and 10 concepts, the set of decimal digits. Recall that LED displays contain 7 light-emitting diodes -- …

classification, generator, multivariate

Computer Science

EITZ Sketch Quality

Humans have used sketching to depict our visual world since prehistoric times. Even today, sketching is possibly the only rendering technique readily avai…

image retrieval, matching, partial, shape retrieval, sketch

Vision

KEGG Metabolic Reaction Netwo…

KEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. In …

classification, clustering, multivariate, regression, text, univariate

Life Sciences

Low Resolution Spectrometer

The Infra-Red Astronomy Satellite (IRAS) was the first attempt to map the full sky at infra-red wavelengths. This could not be done from ground observato…

classification, multivariate

Physical Systems

ADE20k

Scene Parsing Benchmark Scene parsing data and part segmentation data derived from ADE20K dataset could be download from MIT Scene Parsing Benchmark. m…

annotation, benchmark, recognition, scene, segmentation, semantic

Vision

PUT face

9971 images of 100 people

recognition

Vision

Artificial Characters

This database has been artificially generated by using a first order theory which describes the structure of ten capital letters of the English alphabet a…

classification, multivariate

Computer Science

CHALEARN Multi-modal Gesture …

The CHALEARN Multi-modal Gesture Challenge is a dataset +700 sequences for gesture recognition using images, kinect depth, segmentation and skeleton data.…

action, depth, gesture, human, illumination, kinect, recognition, segmentation, skeleton

Vision

The OU-ISIR Gait Database, Tr…

Treadmill gait datasets composed of 34 subjects with 9 speed variations, 68 subjects with 68 subjects, and 185 subjects with various degrees of gait fluct…

recognition

Vision

PAMAP2 Physical Activity Moni…

The PAMAP2 Physical Activity Monitoring dataset contains data of 18 different physical activities (such as walking, cycling, playing soccer, etc.), perfor…

classification, multivariate, time-series

Computer Science

StoneFlakes

Background information: The data set concerns the earliest history of mankind. Prehistoric men created the desired shape of a stone tool by striking on a …

causal-discovery, classification, clustering, multivariate

Others

Page Blocks Classification

The 5473 examples comes from 54 distinct documents. Each observation concerns one block. All attributes are numeric. Data are in a format readable by C4.5.

classification, multivariate

Computer Science

University

Format: Each observation concerns one university. In some cases, more information is provided about the attribute (e.g., units or domain). Some duplicates…

classification, multivariate

Others

Reuters-21578 Text Categoriza…

From the original readme file (please consult it for more information): ------------------------- The documents in the Reuters-21578 collection appeared o…

classification, text

Others

Online News Popularity

* The articles were published by Mashable (www.mashable.com) and their content as the rights to reproduce it belongs to them. Hence, this dataset does not…

classification, multivariate, regression

Business

Molecular Biology (Splice-jun…

Problem Description: Splice junctions are points on a DNA sequence at which `superfluous' DNA is removed during the process of protein creation in highe…

classification, domain-theory, sequential

Life Sciences

Spambase

The "spam" concept is diverse: advertisements for products/web sites, make money fast schemes, chain letters, pornography... Our collection of spam e-mai…

classification, multivariate

Computer Science

Tennis Major Tournament Match…

N/A

classification, clustering, multivariate, regression

Others

Reuter_50_50

The dataset is the subset of RCV1. These corpus has already been used in author identification experiments. In the top 50 authors (with respect to total s…

classification, clustering, domain-theory, multivariate, text

Computer Science

Leaf

For further details on this dataset and/or its attributes, please read the 'ReadMe.pdf' file included and/or consult the Master's Thesis 'Development of a…

classification, multivariate

Computer Science

Flower classification data se…

17 Flower Category Dataset

classification

Vision

Audiology (Standardized)

This database is a standardized version of the original audiology database (see audiology.* in this directory). The non-standard set of attributes have b…

classification, multivariate

Life Sciences

Urban Land Cover

Contains training and testing data for classifying a high resolution aerial image into 9 types of urban land cover. Multi-scale spectral, size, shape, and…

classification, multivariate

Physical Systems

REALDISP Activity Recognition…

The REALDISP (REAListic sensor DISPlacement) dataset has been originally collected to investigate the effects of sensor displacement in the activity recog…

classification, multivariate, time-series

Computer Science

Stanford Background Dataset

The Stanford Background Dataset is a new dataset introduced in Gould et al. (ICCV 2009) for evaluating methods for geometric and semantic scene understand…

classification, geometry, nature, segmentation, semantic, urban

Vision

Connect-4

This database contains all legal 8-ply positions in the game of connect-4 in which neither player has won yet, and in which the next move is not forced. …

classification, multivariate, spatial

Games

KIMA216

The Kimia 216 has 18 classes each consisting of 12 images. It contains shapes silhouettes for birds, bones, brick, camels, car, children, classic cards, e…

animal, binary, kimia, matching, shape retrieval

Vision

Pittsburgh Bridges

There are two versions to the database: - V1 contains the original examples and - V2 contains descriptions after discretizing numeric properti…

classification, multivariate

Others

User Identification From Walk…

The dataset collects data from an Android smartphone positioned in the chest pocket. Accelerometer Data are collected from 22 participants walking in the …

classification, clustering, sequential, time-series, univariate

Others

Nomao

The dataset has been enriched during the Nomao Challenge: [Web Link] organized along with the ALRA workshop (Active Learning in Real-world Applications): …

classification, univariate

Computer Science

Lung Cancer

This data was used by Hong and Young to illustrate the power of the optimal discriminant plane even in ill-posed settings. Applying the KNN method in the …

classification, multivariate

Life Sciences

KUL Belgium Traffic Sign Clas…

BelgiumTSC dataset is built for traffic sign classification purposes. Is is a subset of BelgiumTS dataset and contains cropped images around annotations f…

belgium, classification, road, sign, traffic, urban

Vision

Online Retail

This is a transnational data set which contains all the transactions occurring between 01/12/2010 and 09/12/2011 for a UK-based and registered non-store o…

classification, clustering, multivariate, sequential, time-series

Business

Quality Assessment of Digital…

* The dataset was acquired and annotated by professional physicians at 'Hospital Universitario de Caracas'. * The subjective judgments (target variables) …

classification, multivariate

Life Sciences

Molecular Biology (Promoter G…

This dataset has been developed to help evaluate a "hybrid" learning algorithm ("KBANN") that uses examples to inductively refine preexisting knowledge. …

classification, domain-theory, sequential

Life Sciences

Annealing

N/A

classification, multivariate

Physical Systems

KnapSack

KNAPSACK_01 is a dataset directory which contains some examples of data for 01 Knapsack problems. In the 01 Knapsack problem, we are given a knapsack of…

classification, machine learning

Vision

Dota2 Games Results

Dota 2 is a popular computer game with two teams of 5 players. At the start of the game each player chooses a unique hero with different strengths and wea…

classification, multivariate

Games

Rent3D

The Rent3D dataset comprises floorplans and images. The goal of this work is to enable a 3D virtual-tour of an apartment given a small set of monocular im…

apartment, building, floorplan, indoor, layout, reconstruction, urban

Vision

UIUC Cars

This UIUC Cars dataset by Shivani Agarwal, Aatif Awan and Dan Roth contains images of side views of cars for use in evaluating object detection algorithms…

car, detection, recognition, scale, sideview, urban

Vision

Occupancy Detection

Three data sets are submitted, for training and testing. Ground-truth occupancy was obtained from time stamped pictures that were taken every minute. For …

classification, multivariate, time-series

Computer Science

The OU-ISIR Gait Database, La…

Large population gait datasets composed of 4,016 subjects.

recognition

Vision

YouTube Multiview Video Games…

Please see the README for the details on the data organization, and so on.

classification, clustering, multivariate, text

Computer Science

Chars74K

The Chars74K dataset consists of 64 classes (0-9, A-Z, a-z), 7705 characters obtained from natural images, 3410 hand drawn characters using a tablet PC, 6…

classification, text detection, text recognition

Vision

Text and Vision (TVGraz) Data…

The Text and Vision (TVGraz) dataset is an annotated multi-modal dataset which currently contains 10 visual object categories, 4030 images and associated …

appearance, classification, evaluation, text

Vision

PASCAL VOC 2009 dataset

Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets

classification, detection

Vision

EMG Physical Action Data Set

1. Protocol: Three male and one female subjects (age 25 to 30), who have experienced aggression in scenarios such as physical fighting, took part in…

classification, time-series

Physical Systems

ASL Datasets Repository

This site is dedicated to provide datasets for the Robotics community with the aim to facilitate result evaluations and comparisons. The datasets presente…

3d, city, laser, nature, urban

Vision

Lymphography

This is one of three domains provided by the Oncology Institute that has repeatedly appeared in the machine learning literature. (See also breast-cancer a…

classification, multivariate

Life Sciences

Pima Indians Diabetes

Several constraints were placed on the selection of these instances from a larger database. In particular, all patients here are females at least 21 year…

classification, multivariate

Life Sciences

Mushroom

This data set includes descriptions of hypothetical samples corresponding to 23 species of gilled mushrooms in the Agaricus and Lepiota Family (pp. 500-52…

classification, multivariate

Life Sciences

THUR15000

We introduce a labeled dataset of categorized images for evaluating sketch based image retrieval. Using Flickr, we downloaded about 3000 images for each o…

attention, group, internet, retrieval, saliency, salient object detection, shape, sketch, visual

Vision

Qualitative_Bankruptcy

The parameters which we used for collecting the dataset is referred from the paper 'The discovery of experts decision rules from qualitative bankruptcy da…

classification, multivariate

Computer Science

EEG Eye State

All data is from one continuous EEG measurement with the Emotiv EEG Neuroheadset. The duration of the measurement was 117 seconds. The eye state was detec…

classification, multivariate, sequential, time-series

Life Sciences

xawAR16

The xawAR16 dataset is a multi-RGBD camera dataset, generated inside an operating room (IHU Strasbourg), which was designed to evaluate tracking/relocaliz…

depth, medicine, operation, recognition, surgery, table, video

Vision

Swedish Traffic Sign Recognit…

The Swedish Traffic Sign Recognition provides Matlab code for parsing the annotation files and displaying the results. Part0 for each set contains the ann…

city, detection, recognition, sign, traffic, urban

Vision

Mammographic Mass

Mammography is the most effective method for breast cancer screening available today. However, the low positive predictive value of breast biopsy resultin…

classification, multivariate

Life Sciences

Gas sensor array under dynami…

This data set contains the acquired time series from 16 chemical sensors exposed to gas mixtures at varying concentration levels. In particular, we genera…

classification, multivariate, regression, time-series

Computer Science

Hayes-Roth

This database contains 5 numeric-valued attributes. Only a subset of 3 are used during testing (the latter 3). Furthermore, only 2 of the 3 concepts are…

classification, multivariate

Social Sciences

Northix

Northix is designed to be a schema matching benchmark problem for data integration of two entity relationship databases. Northix is the resulting schema m…

classification, multivariate, text, univariate

Computer Science

PubFig: Public Figures Face D…

The PubFig database is a large, real-world face dataset consisting of 58,797 images of 200 people collected from the internet. Unlike most other existing …

recognition

Vision

Contraceptive Method Choice

This dataset is a subset of the 1987 National Indonesia Contraceptive Prevalence Survey. The samples are married women who were either not pregnant or do …

classification, multivariate

Life Sciences

Educational Process Mining (E…

The experiments have been carried out with a group of 115 students of first-year, undergraduate Engineering major of the University of Genoa. We carried…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

Statlog (Shuttle)

Approximately 80% of the data belongs to class 1. Therefore the default accuracy is about 80%. The aim here is to obtain an accuracy of 99 - 99.9%. The e…

classification, multivariate

Physical Systems

Motion Capture Hand Postures

A Vicon motion capture camera system was used to record 12 users performing 5 hand postures with markers attached to a left-handed glove. A rigid patter…

classification, clustering, multivariate

Computer Science

SUSY

Provide all relevant informatioThe data has been produced using Monte Carlo simulations. The first 8 features are kinematic properties measured by the par…

classification

Physical Systems

MHEALTH Dataset

The MHEALTH (Mobile HEALTH) dataset comprises body motion and vital signs recordings for ten volunteers of diverse profile while performing several physic…

classification, multivariate, time-series

Computer Science

Covertype

Predicting forest cover type from cartographic variables only (no remotely sensed data). The actual forest cover type for a given observation (30 x 30 me…

classification, multivariate

Life Sciences

Las Vegas Strip

All the 504 reviews were collected between January and August of 2015.

classification, regression

Business

ICG PRID 2011

The Person Re-ID (PRID) 2011 dataset was created in co-operation with the Austrian Institute of Technology for the purpose of testing person re-identifica…

appearance, change, classification, graz, identification, illumination, multiview, pedestrian, trajectory

Vision

Balance Scale

This data set was generated to model psychological experimental results. Each example is classified as having the balance scale tip to the right, tip to …

classification, multivariate

Social Sciences

Multi-Camera Action Dataset

An indoor action recognition dataset which consists of 18 classes performed by 20 individuals. Each action is individually performed for 8 times (4 daytim…

action, cross-view, indoor, multi-camera, open-view, recognition, video

Vision

TUD Shapes 1+2

This material is supplementary to Michael Stark, Bernt Schiele. How Good are Local Features for Classes of Geometric Objects. Eleventh IEEE Internatio…

binary, classification, object, shape, tool

Vision

Shefeld Kinect Gesture (SKIG)…

The Shefeld Kinect Gesture (SKIG) dataset contains 2160 hand gesture sequences (1080 RGB sequences and 1080 depth sequences) collected from 6 subjects. Al…

action, depth, gesture, human, illumination, kinect, recognition

Vision

Statlog (Heart)

Cost Matrix _______ abse pres absence 0 1 presence 5 0 where the rows represent the true values and the columns the predicted.

classification, multivariate

Life Sciences

KIMIA25

The Kimia 25 consists of 6 classes and 25 images. They are part of the Shape Indexing of Image Database (SIID) project, which also contains the SIID silho…

binary, kimia, matching, shape retrieval

Vision

Daimler Mono Pedestrian Class…

The Daimler Mono Pedestrian Classification Benchmark dataset consists of two parts: a base data set. The base data set contains a total of 4000 pedestri…

classification, illumination, object, outdoor, pedestrian, scale, urban

Vision

JPL First-Person Interaction

JPL First-Person Interaction dataset (JPL-Interaction dataset) is composed of human activity videos taken from a first-person viewpoint. The dataset parti…

action, human, interactive, motion, recognition, video

Vision

Our Database of Faces

The Our Database of Faces (ORL) dataset contains ten different images of each of 40 distinct subjects. For some subjects, the images were taken at differe…

expression, face, human, illumination, recognition

Vision

ICG Annotated Facial Landmark…

The Annotated Facial Landmarks in the Wild (AFLW) consists of a large-scale collection of annotated face images gathered from the web, exhibiting a large …

age, annotation, detection, face, landmark, pose

Vision

microblogPCU

Our dataset is used by us to explore spammers in microblog and you can access our demo system at [Web Link]Please add :8080 after the domain name as port.…

causal-discovery, classification, multivariate, sequential, text, univariate

Computer Science

Cuff-Less Blood Pressure Esti…

The main goal of this data set is providing clean and valid signals for designing cuff-less blood pressure estimation algorithms. The raw electrocardiogra…

classification, multivariate, regression

Life Sciences

Buzz in social media

Please see [Web Link]

classification, multivariate, regression, time-series

Computer Science

Crowdsourced Mapping

This dataset was derived from geospatial data from two sources: 1) Landsat time-series satellite imagery from the years 2014-2015, and 2) crowdsourced geo…

classification, multivariate

Physical Systems

FaceScrub Face Dataset

The FaceScrub dataset is a real-world face dataset comprising 107,818 face images of 530 male and female celebrities detected in images retrieved from the…

recognition

Vision

MSRC-12: Kinect gesture data …

The Microsoft Research Cambridge-12 Kinect gesture data set consists of sequences of human movements, represented as body-part locations, and the associat…

recognition

Vision

Connectionist Bench (Sonar, M…

The file "sonar.mines" contains 111 patterns obtained by bouncing sonar signals off a metal cylinder at various angles and under various conditions. The …

classification, multivariate

Physical Systems

Dresses_Attribute_Sales

Style, Price, Rating, Size, Season, NeckLine, SleeveLength, waiseline, Material, FabricType, Decoration, Pattern, Type, Recommendation are Attributes in d…

classification, clustering, text

Computer Science

Statlog (Vehicle Silhouettes)

The purpose is to classify a given silhouette as one of four types of vehicle, using a set of features extracted from the silhouette. The vehicle may be …

classification, multivariate

Others

ILPD (Indian Liver Patient Da…

This data set contains 416 liver patient records and 167 non liver patient records.The data set was collected from north east of Andhra Pradesh, India. Se…

classification, multivariate

Life Sciences

Smartphone Dataset for Human …

This dataset is an addition to the dataset at [Web Link] We collected more dataset to improve the accuracy of our HAR algorithms applied in …

classification, time-series

Computer Science

DBWorld e-mails

I collected 64 e-mails from DBWorld newsletter and I used them to train different algorithms in order to classify between 'announces of conferences' and '…

classification, text

Computer Science

Spoken Arabic Digit

Dataset from 8800(10 digits x 10 repetitions x 88 speakers) time series of 13 Frequency Cepstral Coefficients (MFCCs) had taken from 44 males and 44 femal…

classification, multivariate, time-series

Others

Detail 2D Projection DataSet

Detail 2D Projection DataSet is a database of 2d projections of mechanical details with holes. The dataset consists of 13 shape categories where each cate…

binary, detail, holes, matching, shape retrieval

Vision

POSTECH Labeled Faces in the …

POS Labeled Faces in the Wild, a collection of face which is proposed for studying face identification in unconstrained environment, its purpose is servin…

face, identification, recognition, registration, wild

Vision

image panorama gdbicp

Generalized Dual Bootstrap-ICP Algorithm

image registration, matching, panorama

Vision

Australian Sign Language sign…

Data was captured using a setup that consisted of: - Two Fifth Dimension Technologies (5DT) gloves, one right and one left - Two Ascension Flock-of-Bir…

classification, multivariate, time-series

Others

MicroMass

This MALDI-TOF dataset consists in:A) A reference panel of 20 Gram positive and negative bacterial species covering 9 genera among which several species a…

classification, multivariate

Life Sciences

Flags

This data file contains details of various nations and their flags. In this file the fields are separated by spaces (not commas). With this data you can …

classification, multivariate

Others

KEGG Metabolic Relation Netwo…

KEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. In …

classification, clustering, multivariate, regression, text, univariate

Life Sciences

p53 Mutants

Biophysical models of mutant p53 proteins yield features which can be used to predict p53 transcriptional activity. All class labels are determined via i…

classification, multivariate

Life Sciences

GeoFaces

A large dataset of geotagged face images collected from Flickr. The zip file contains text files containing urls of the images. Face2GPS: Estimating Geo…

age, classification, face, gender, geotagged, human, localization

Vision

MONK's Problems

The MONK's problem were the basis of a first international comparison of learning algorithms. The result of this comparison is summarized in "The MONK's P…

classification, multivariate

Others

Dubrovnik6K and Rome16K

The Dubrovnik6K and Rome16K datasets are image collections for SfM reconstruction, where the suffix refers to the number of images in the dataset. Dubro…

3d reconstruction, dubrovnik, landmark, rome, sfm, urban

Vision

Forest type mapping

This data set contains training and testing data from a remote sensing study which mapped different forest types based on their spectral characteristics a…

classification, multivariate

Life Sciences

Multispectral Imaging (MSI)

Multispectral Imaging (MSI) datasets were acquired using IRIS II which is a lightweight portable system comprising of a high resolution camera, a novel fi…

alignment, groundtruth, illumination, matching, multi-spectral, registration, wavelength

Vision

Open Images Dataset

Today, we introduce Open Images, a dataset consisting of ~9 million URLs to images that have been annotated with labels spanning over 6000 categories. We …

annotation, automatic, category, classification, deep, image, large-scale, real

Vision

HIGGS

The data has been produced using Monte Carlo simulations. The first 21 features (columns 2-22) are kinematic properties measured by the particle detectors…

classification

Physical Systems

Statlog (Image Segmentation)

The instances were drawn randomly from a database of 7 outdoor images. The images were handsegmented to create a classification for every pixel. Each …

classification, multivariate

Others

Rijksmuseum Challenge Dataset…

Over 110,000 photographic reproductions of the artworks exhibited in the Rijksmuseum (Amsterdam, the Netherlands). Offers four automatic visual recognitio…

recognition

Vision

URL Reputation

Uncompressing the archive url_svmlight.tar.gz will yield a directory url_svmlight/ containing the following files: * FeatureTypes --- A text file list…

classification, multivariate, time-series

Computer Science

First-order theorem proving

See the file bridge-holden-paulson-details.txt in the submitted tarball.

classification, multivariate

Computer Science

MAGIC Gamma Telescope

The data are MC generated (see below) to simulate registration of high energy gamma particles in a ground-based atmospheric Cherenkov gamma telescope usin…

classification, multivariate

Physical Systems

Gas Sensor Array Drift Dataset

This archive contains 13910 measurements from 16 chemical sensors utilized in simulations for drift compensation in a discrimination task of 6 gases at va…

classification, multivariate

Computer Science

MiniBooNE particle identifica…

The submitted file is set up as follows. In the first line is the the number of signal events followed by the number of background events. The signal even…

classification, multivariate

Physical Systems

Diabetes 130-US hospitals for…

The dataset represents 10 years (1999-2008) of clinical care at 130 US hospitals and integrated delivery networks. It includes over 50 features representi…

classification, clustering, multivariate

Life Sciences

Samantha

The SAMANTHA (Structure-and-Motion Pipeline on a Hierarchical Cluster Tree) dataset contains 4 sequences for 3D reconstruction: Pozzoveggiani, Piazza Dant…

3d reconstruction, geometry, landmark, model fitting, sfm

Vision

LSVT Voice Rehabilitation

The original paper demonstrated that it is possible to correctly replicate the experts' binary assessment with approximately 90% accuracy using both 10-fo…

classification, multivariate

Life Sciences

Statlog (Landsat Satellite)

The database consists of the multi-spectral values of pixels in 3x3 neighbourhoods in a satellite image, and the classification associated with the centra…

classification, multivariate

Physical Systems

ChokePoint Dataset

ChokePoint is a video dataset designed for experiments in person identification/verification under real-world surveillance conditions. The dataset consist…

recognition

Vision

Synthetic Control Chart Time …

This dataset contains 600 examples of control charts synthetically generated by the process in Alcock and Manolopoulos (1999). There are six different cla…

classification, clustering, time-series

Others

MPEG-7 Core Experiment CE-Sha…

MPEG-7 Core Experiment CE-Shape-1 [?] is a popular database for shape matching evaluation consisting of 70 shape categories, where each category is repres…

binary, bullseye, matching, shape retrieval

Vision

Climate Model Simulation Cras…

This dataset contains records of simulation crashes encountered during climate model uncertainty quantification (UQ) ensembles. Ensemble members were co…

classification, multivariate

Physical Systems

SIPI textures

The Textures volume currently contains 154 images, all monochrome, 129 512x512 and 25 1024x1024. For the Brodatz texture images, the number in parenthes…

benchmark, classification, evaluation, segmentation, synthetic, texture

Vision

Dataset for Sensorless Drive …

Features are extracted from electric current drive signals. The drive has intact and defective components. This results in 11 different classes with diffe…

classification, multivariate

Computer Science

Australian Sign Language signs

The source of the data is the raw measurements from a Nintendo PowerGlove. It was interfaced through a PowerGlove Serial Interface to a Silicon Graphics 4…

classification, multivariate, time-series

Others

MSR RGB-D 7-Scenes

The MSR RGB-D Dataset 7-Scenes dataset is a collection of tracked RGB-D camera frames. The dataset may be used for evaluation of methods for different app…

depth, kinect, location, reconstruction, tracking, video

Vision

UJI Pen Characters

We create a character database by collecting samples from 11 writers. Each writer contributed with letters (lower and uppercase), digits, and other chara…

classification, multivariate, sequential

Computer Science

Dataset for ADL Recognition w…

The Dataset for ADL Recognition with Wrist-worn Accelerometer is a public collection of labelled accelerometer data recordings to be used for the creation…

classification, clustering, multivariate, time-series

Computer Science

PubChem Bioassay Data

21 bioassay datasets generated from Pubchem. Both Primary and confirmatory bioassays (12 bioassays, 21 mixes)The data is provided in the same train/test s…

classification, multivariate

Life Sciences

YorkUrbanDB

The York Urban Line Segment Database is a compilation of 102 images (45 indoor, 57 outdoor) of urban environments consisting mostly of scenes from the cam…

geometry, manhattan, outdoor, pose estimation, reconstruction, urban, vanishing point

Vision

Poker Hand

Each record is an example of a hand consisting of five playing cards drawn from a standard deck of 52. Each card is described using two attributes (suit a…

classification, multivariate

Games

Chess (King-Rook vs. King-Kni…

The companion file is a Common Lisp demonstration file that generates knight-pin Chess end-game samples. Start up Lisp and load the file. It generates 10…

classification, generator, multivariate

Games

SHREC

Unlike the previous SHREC contests, the objective of this SHREC 2012 contest is to evaluate the performance of 3D-mesh segmentation techniques instead of …

3d, mesh, part, segmentation

Vision

Car Evaluation

Car Evaluation Database was derived from a simple hierarchical decision model originally developed for the demonstration of DEX, M. Bohanec, V. Rajkovic: …

classification, multivariate

Others

Wall-Following Robot Navigati…

The provided files comprise three different data sets. The first one contains the raw values of the measurements of all 24 ultrasound sensors and the cor…

classification, multivariate, sequential

Computer Science

Hybrid Indoor Positioning Dat…

The measurements were created to ease the development, comparison and evaluation of fingerprinting based hybrid indoor positioning methods. The measuremen…

classification, multivariate, sequential, time-series

Computer Science

YouTube Faces

The data set contains 3,425 videos of 1,595 different people. The shortest clip duration is 48 frames, the longest clip is 6,070 frames, and the average l…

recognition

Vision

Amazon Commerce reviews set

dataset are derived from the customers reviews in Amazon Commerce Website for authorship identification. Most previous studies conducted the identifi…

classification, domain-theory, multivariate, text

Physical Systems

UJIIndoorLoc-Mag

Indoor localization is a key topic for mobile computing. However, it is still very difficult for the mobile sensing community to compare state-of-art Indo…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

Parkinson Disease Spiral Draw…

The PD and control handwriting database consists of 62 PWP (People with parkinson) and 15 healthy individuals who appealed at the Department of Neurology …

classification, clustering, multivariate, regression

Computer Science

ser Knowledge Modeling Data (…

-- The users' knowledge class were classified by the authors using intuitive knowledge classifier (a hybrid ML technique of k-NN and meta-heuristic …

classification, multivariate

Computer Science

Geo-Magnetic field and WLAN d…

Indoor localisation is a key topic for the Ambient Intelligence (AmI) research community. In this scenarios, recent advancements in wearable technologie…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

City planar and non-planar

The city planar and non-planar datset consists of urban scenes accompanied by text files describing the plane/non-plane locations. Training Set (Univer…

3d, building, detection, estimation, plane, urban

Vision

Aachen Retrieval

The Aachen dataset consists of 4479 images taken with multiple cameras (3GB), 369 query images taken with the camera of a mobile phone together with their…

3d reconstruction, aachen, image retrieval, landmark, sfm

Vision

Bristol Egocentric Object Int…

The BEOID dataset includes object interactions ranging from preparing a coffee to operating a weight lifting machine and opening a door. The dataset is re…

3d, egocentric, interaction, object, pose, tracking, video

Vision

Dow Jones Index

In predicting stock prices you collect data over some period of time - day, week, month, etc. But you cannot take advantage of data from a time period unt…

classification, clustering, time-series

Business

California-ND

An Annotated Dataset For Near-Duplicate Detection In Personal Photo Collections Managing photo collections involves a variety of image quality assessmen…

copyright, detection, duplicate, groundtruth, retrieval

Vision

KDC-4007 dataset Collection

The most important feature of this dataset is its simplicity to use and its being well-documented, which can be widely used in various studies of text ana…

classification, multivariate, regression, text

Computer Science

ETHZ Shape Classes

A dataset for testing object class detection algorithms. It contains 255 test images and features five diverse shape-based classes (apple logos, bottles, …

classification

Vision

Street View Text

The Street View Text (SVT) dataset contains 647 words and 3796 letters in 249 images harvested from Google Street View. The dataset is more challenging …

classification, outdoor, text detection, text recognition, urban

Vision

Caltech Buildings Dataset

The Caltech Buildings dataset consists of images taken for 50 buildings around the Caltech campus. Five different images were taken for each building from…

building, caltech, hierarchy, retrieval, taxonomy, urban

Vision

sEMG for Basic Hand movements

Instrumentation: The data were collected at a sampling rate of 500 Hz, using as a programming kernel the National Instruments (NI) Labview. The signals we…

classification, time-series

Life Sciences

Demospongiae

This dataset contains 503 sponges belonging to the Demospongiae class collected from the Mediterranean (451 sponges) and Atlantic oceans (52 sponges). Eac…

classification, multivariate

Life Sciences

Chronic_Kidney_Disease

We use the following representation to collect the dataset age - age bp - blood pressure sg - specific gravity al - …

classification, multivariate

Others

Japanese Vowels

The data was collected for examining our newly developed classifier for multidimensional curves (multidimensional time series). Nine male speakers uttered…

classification, multivariate, time-series

Others

Bach Choral Harmony

Pitch classes information has been extracted from MIDI sources downloaded from (JSB Chorales)[[Web Link]]. Meter information has been computed through the…

classification, sequential

Others

Soybean (Small)

A small subset of the original soybean database. See the reference for Fisher and Schlimmer in soybean-large.names for more information. Steven Souders …

classification, multivariate

Life Sciences

Thoracic Surgery Data

The data was collected retrospectively at Wroclaw Thoracic Surgery Centre for patients who underwent major lung resections for primary lung cancer in the …

classification, multivariate

Life Sciences

ChairGest Gestures

ChairGest is an open challenge / benchmark. The task consists in spotting and recognizing gestures from multiple synchronized sensors: 1 Kinect and 4 Xse…

benchmark, detection, gesture, human, kinect, recognition

Vision

Musk (Version 1)

This dataset describes a set of 92 molecules of which 47 are judged by human experts to be musks and the remaining 45 molecules are judged to be non-musks…

classification, multivariate

Physical Systems

default of credit card clients

This research aimed at the case of customers default payments in Taiwan and compares the predictive accuracy of probability of default among six data mini…

classification, multivariate

Business

Energy efficiency

We perform energy analysis using 12 different building shapes simulated in Ecotect. The buildings differ with respect to the glazing area, the glazing are…

classification, multivariate, regression

Computer Science

Libras Movement

The dataset (movement_libras) contains 15 classes of 24 instances each, where each class references to a hand movement type in LIBRAS. In the video pre-p…

classification, clustering, multivariate, sequential

Others

Stanford 40 Actions

The Stanford 40 Actions dataset contains images of humans performing 40 actions. In each image, we provide a bounding box of the person who is performing …

action, boundingbox, detection, human, recognition

Vision

Paris Retrieval

The Paris dataset consists of 6412 images. Images have high resolution and are in JPEG format. http://www.robots.ox.ac.uk/~vgg/data/parisbuildings/pari…

image retrieval, landmark, paris, urban

Vision

Vertebral Column

Biomedical data set built by Dr. Henrique da Mota during a medical residence period in the Group of Applied Research in Orthopaedics (GARO) of the Centre …

classification, multivariate

Others

Feret

Face and Gesture Recognition Working Group FGnet

facial, recognition

Vision

Farm Ads

This data was collected from text ads found on twelve websites that deal with various farm animal related topics. Information from the ad creative and th…

classification, text

Business

Activity Recognition from Sin…

--- The dataset collects data from a wearable accelerometer mounted on the chest --- Sampling frequency of the accelerometer: 52 Hz --- Accelerom…

classification, clustering, sequential, time-series, univariate

Others

Face and Gesture Recognition …

Face and Gesture Recognition Working Group FGnet

recognition

Vision

Video classification USAA dat…

The USAA dataset includes 8 different semantic class videos which are home videos of social occassions which feature activities of group of people. It con…

classification

Vision

Congressional Voting Records

This data set includes votes for each of the U.S. House of Representatives Congressmen on the 16 key votes identified by the CQA. The CQA lists nine diff…

classification, multivariate

Social Sciences

Diabetic Retinopathy Debrecen…

This dataset contains features extracted from the Messidor image set to predict whether an image contains signs of diabetic retinopathy or not. All featur…

classification, multivariate

Life Sciences

News Aggregator

News are grouped into clusters that represent pages discussing the same news story. The dataset includes also references to web pages that, at the access…

classification, clustering, multivariate

Others

EITZ Sketch-Based Image Retri…

We introduce a benchmark for evaluating the performance of large scale sketch-based image retrieval systems. The necessary data is acquired in a controlle…

image retrieval, matching, partial, shape retrieval, sketch

Vision

Perfume Data

The data set gathered when we were working at project for Bahrain university between 2002 and 2003.

classification, clustering, domain-theory, univariate

Computer Science

Drug consumption (quantified)

Database contains records for 1885 respondents. For each respondent 12 attributes are known: Personality measurements which include NEO-FFI-R (neuroticism…

classification, multivariate

Social Sciences

TV News Channel Commercial De…

Automatic identification of commercial blocks in news videos finds a lot of applications in the domain of television broadcast analysis and monitoring. Co…

classification, clustering, multivariate

Computer Science

TTC-3600: Benchmark dataset f…

The dataset consists of a total of 3600 documents including 600 news/texts from six categories economy, culture-arts, health, politics, sports and techno…

classification, clustering, text

Computer Science

Tamilnadu Electricity Board H…

Collect the real time readings for residential,commercial,industrial,agriculure,to find the accuracy consumption in Tamil Nadu Around Thanajvur

classification, clustering, multivariate, regression

Life Sciences

Parkinsons

This dataset is composed of a range of biomedical voice measurements from 31 people, 23 with Parkinson's disease (PD). Each column in the table is a parti…

classification, multivariate

Life Sciences

Smartphone-Based Recognition …

The experiments were carried out with a group of 30 volunteers within an age bracket of 19-48 years. They performed a protocol of activities composed of s…

classification, multivariate, time-series

Life Sciences

Stroke Width Transform Text

Stroke Width Transform Text dataset is by Boris Epstein and consists of 307 images and XXX text instances. Detecting Text in Natural Scenes with Stroke…

classification, text detection, text recognition

Vision

Localization Data for Person …

People used for recording of the data were wearing four tags (ankle left, ankle right, belt and chest). Each instance is a localization data for one of t…

classification, sequential, time-series, univariate

Life Sciences

Ian Dworkin (McMaster Univers…

This is the database of biological images (from the genetics model system, Drosophila melanogaster, a fruit fly) across multiple levels of variation. we…

animal, biology, classification, fly, genetic, variation

Vision

Semeion Handwritten Digit

1593 handwritten digits from around 80 persons were scanned, stretched in a rectangular box 16x16 in a gray scale of 256 values.Then each pixel of each im…

classification, multivariate

Computer Science

Ecoli

The references below describe a predecessor to this dataset and its development. They also give results (not cross-validated) for classification by a rule…

classification, multivariate

Life Sciences

Badges

Part of the problem in using an automated program to discover the unknown target function is to decide how to encode names such that the program can be us…

classification, text, univariate

Others

Vicon Physical Action Data Set

1. Protocol: Seven male and three female subjects (age 25 to 30), who have experienced aggression in scenarios such as physical fighting, took part …

classification, time-series

Physical Systems

Outex texture bench

The Outex dataset is part of a framework for empirical evaluation of texture classification and segmentation algorithms. The framework is being construct…

benchmark, classification, segmentation, synthetic, texture

Vision

Mechanical Analysis

F. Bergadano supplied this database. Each instance contains many components, each of which has 8 attributes. Different instances in this database have d…

classification, multivariate

Computer Science

PASCAL Context

We would like to announce the release of PASCAL-Context dataset. We augmented PASCAL VOC 2010 dataset with annotations for 400+ additional categories. In …

benchmark, category, dense, pascal, recognition, segmentation, semantic, shape

Vision

DrivFace

The DrivFace database contains images sequences of subjects while driving in real scenarios. It is composed of 606 samples of 640480 pixels each, acquired…

classification, clustering, multivariate, regression

Computer Science

Epileptic Seizure Recognition

Please find the original data at '[Web Link]'

classification, clustering, multivariate, time-series

Life Sciences

CNAE-9

This is a data set containing 1080 documents of free text business descriptions of Brazilian companies categorized into a subset of 9 categories cataloged…

classification, multivariate, text

Business

Dexter

The original data were formatted by Thorsten Joachims in the bag-of-words representation. There were 9947 features (of which 2562 are always zeros for all…

classification, multivariate

Others

QMUL underGround Re-IDentific…

This dataset contains 250 pedestrian image pairs + 775 additional images captured in a busy underground station for the research on person re-identificati…

recognition

Vision

Statlog (German Credit Data)

Two datasets are provided. the original dataset, in the form provided by Prof. Hofmann, contains categorical/symbolic attributes and is in the file "germ…

classification, multivariate

Financial

Labeled Faces in the Wild

A database of face photographs designed for studying the problem of unconstrained face recognition

recognition

Vision

Cardiotocography

2126 fetal cardiotocograms (CTGs) were automatically processed and the respective diagnostic features measured. The CTGs were also classified by three exp…

classification, multivariate

Life Sciences

Happy People Images Database

Group emotion recognition in images - Happiness Intensity labels for group of people in images. The images have been collected from Flickr using keyword s…

behavior, emotion, facial expression, flickr, group, human, wild

Vision

action recognition benchmark

We wanted to have a collection of action recognition papers and results that everybody can use for reference. The site will work by the community principl…

action, benchmark, dataset, recognition

Vision

Human Activity Recognition Us…

The experiments have been carried out with a group of 30 volunteers within an age bracket of 19-48 years. Each person performed six activities (WALKING, W…

classification, clustering, multivariate, time-series

Computer Science

McGill Real-World Face Video …

This database contains 18000 video frames of 640x480 resolution from 60 video sequences, each of which recorded from a different subject (31 female and 29…

classification

Vision

Geographical Original of Music

The dataset was built from a personal collection of 1059 tracks covering 33 countries/area. The music used is traditional, ethnic or `world' only, as cla…

classification, multivariate, regression

Others

Indoor User Movement Predicti…

This dataset represents a real-life benchmark in the area of Ambient Assisted Living applications, as described in [1]. The binary classification task con…

classification, multivariate, sequential, time-series

Computer Science

MEU-Mobile KSD

The dataset is used in the evaluation of EER, FRR and FAR metrics using a new anomaly detector model (Med-Min-Diff). The typed text in the experiment is t…

classification, multivariate

Computer Science

Open University Learning Anal…

Open University Learning Analytics Dataset (OULAD) contains data about courses, students and their interactions with Virtual Learning Environment (VLE) fo…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

All I Have Seen (AIHS)

The All I Have Seen (AIHS) dataset is created to study the properties of total visual input in humans, for around two weeks Nebojsa Jojic wore a camera ca…

3d, clustering, indoor, outdoor, scene, similarity, study, summary, user, video

Vision

Credit Approval

This file concerns credit card applications. All attribute names and values have been changed to meaningless symbols to protect confidentiality of the da…

classification, multivariate

Financial

Japanese Credit Screening

Examples represent positive and negative instances of people who were and were not granted credit. The theory was generated by talking to the individual…

classification, domain-theory, multivariate

Financial

Machine Learning based ZZAlph…

1. Title of Database: Machine Learning based ZZAlpha Stock Recommendations 2. Sources: (a) Original owners of data: ZZAlpha Ltd., 4729 E. Sunrise #109,…

classification, sequential, time-series

Business

Hepatitis

Please ask Gail Gong for further information on this database.

classification, multivariate

Life Sciences

Cylinder Bands

Here's the abstract from the above reference: ABSTRACT: Machine learning tools show significant promise for knowledge acquisition, particularly when huma…

classification, multivariate

Physical Systems

Ionosphere

This radar data was collected by a system in Goose Bay, Labrador. This system consists of a phased array of 16 high-frequency antennas with a total trans…

classification, multivariate

Physical Systems

Glass Identification

Vina conducted a comparison test of her rule-based system, BEAGLE, the nearest-neighbor algorithm, and discriminant analysis. BEAGLE is a product availab…

classification, multivariate

Physical Systems

HandNet annotated hand dataset

The HandNet dataset contains depth images of 10 participants hands non-rigidly deforming infront of a RealSense RGB-D camera. This dataset includes 214…

articulation, classification, detection, fingertip, hand, pose, rgbd, segmentation, video

Vision

Meta-data

This DataSet is about the results of Statlog project. The project performed a comparative study between Statistical, Neural and Symbolic learning algorith…

classification, multivariate

Others

PASCAL VOC Parts

The PASCAL VOC is augmented with segmentation annotation for semantic parts of objects. For example, for the person category, we provide segmentation mask…

detection, human, object, part, pascal, pedestrian, recognition, segmentation, semantic

Vision

Lenses

The examples are complete and noise free. The examples highly simplified the problem. The attributes do not fully describe all the factors affecting the d…

classification, multivariate

Others

Fertility

Provide all relevant information about your data set.

classification, multivariate, regression

Life Sciences

Gisette

The digits have been size-normalized and centered in a fixed-size image of dimension 28x28. The original data were modified for the purpose of the feature…

classification, multivariate

Computer Science

HTRU2

HTRU2 is a data set which describes a sample of pulsar candidates collected during the High Time Resolution Universe Survey (South) [1]. Pulsars are a r…

classification, clustering, multivariate

Physical Systems

SPECTF Heart

The dataset describes diagnosing of cardiac Single Proton Emission Computed Tomography (SPECT) images. Each of the patients is classified into two categor…

classification, multivariate

Life Sciences

Data for Software Engineering…

The data can be used to try to predict student learning in SE teamwork based on observation of their team activity **** README FILE from the submitted d…

classification, sequential, time-series

Computer Science

Musk (Version 2)

This dataset describes a set of 102 molecules of which 39 are judged by human experts to be musks and the remaining 63 molecules are judged to be non-musk…

classification, multivariate

Physical Systems

e-Lab Video Data Set

Video data sets to train machines to recognise objects in our environment. e-VDS35 has 35 classes and a total of 2050 videos of roughly 10 seconds each.

classification

Vision

Omnidirectional and panoramic…

We share our omnidirectional and panoramic image dataset (with annotations) to be used for human and car detection. Please reach through: http://cvrg.iyt…

car, detection, human, omnidirection, panorama, recognition

Vision

UJI Pen Characters (Version 2)

We have created the UJIpenchars2 character database by collecting samples from 60 writers at two different sites in two phases: 1st phase, 11 writers, car…

classification, multivariate, sequential

Computer Science

AutoUniv

The user first creates a classification model and then generates classified examples from it. To create a model, the following are specified: the number o…

classification, multivariate

Others

Folio

- The leaves were placed on a white background and then photographed. - The pictures were taken in broad daylight to ensure optimum light intensity.

classification, clustering, multivariate

Others

Person identification in TV s…

Face tracks, features and shot boundaries from our latest CVPR 2013 paper. It is obtained from 6 episodes of Buffy the Vampire Slayer and 6 episodes of Bi…

recognition

Vision

QSAR biodegradation

The QSAR biodegradation dataset was built in the Milano Chemometrics and QSAR Research Group (Universit degli Studi Milano Bicocca, Milano, Italy). The r…

classification, multivariate

Others

LabelMeFacade

The LabelMeFacade dataset contains buildings, windows, sky and a limited number of unlabeled regions (maximally 20% covering of the image). This procedure…

facade, recognition, rectified, segmentation, semantic, urban

Vision

Shuttle Landing Control

This is a tiny database. Michie reports that Burke's group used RULEMASTER to generate comprehendable rules for determining the conditions under which an…

classification, multivariate

Physical Systems

NYU Depth v2

The NYU-Depth V2 data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft…

depth, kinect, label, reconstruction, semantic segmentation

Vision

Cholec80

The Cholec80 dataset contains 80 videos of cholecystectomy surgeries performed by 13 surgeons. The videos are captured at 25 fps. The dataset is labeled w…

medicine, phase, recognition, surgery, tool, video

Vision

YouTube Comedy Slam Preferenc…

YouTube Comedy Slam ([Web Link]) is a video discovery experiment running on YouTube's version of labs (called TestTube) for a few months in 2011 and 2012.…

classification, text

Computer Science

Gas sensor array under flow m…

The measured data was collected using a chemical sensing system based on an array of 16 metal-oxide gas sensors and an external mechanical ventilator to s…

classification, multivariate, regression, time-series

Computer Science

One-hundred plant species lea…

For Each feature, a 64 element vector is given per sample of leaf. These vectors are taken as a contigous descriptors (for shape) or histograms (for textu…

classification

Life Sciences

Primary Tumor

This is one of three domains provided by the Oncology Institutenthat has repeatedly appeared in the machine learning literature. (See also breast-cancer …

classification, multivariate

Life Sciences

ICDAR 2003

The ICDAR 2003 datasets available for download on this site: Robust Reading , Robust Word Recognition , Robust OCR , Text Locating and Cursive Script . …

classification, text detection, text recognition

Vision

Hieroglyph Dataset

Ancient Egyptian Hieroglyph Dataset.

recognition

Vision

Oxford Buildings

The Oxford Buildings dataset by James Philbin and Andrew Zisserman consists of 5062 images collected from Flickr by searching for particular Oxford landma…

image retrieval, landmark, oxford, urban

Vision

Skin Segmentation

The skin dataset is collected by randomly sampling B,G,R values from face images of various age groups (young, middle, and old), race groups (white, black…

classification, univariate

Computer Science

Gas sensors for home activity…

This dataset has recordings of a gas sensor array composed of 8 MOX gas sensors, and a temperature and humidity sensor. This sensor array was exposed to b…

classification, multivariate, time-series

Computer Science

Wholesale customers

Provide all relevant information about your data set.

classification, clustering, multivariate

Business

Miskolc IIS Hybrid IPS

The measurements were created to ease the development, comparison and evaluation of fingerprinting based hybrid indoor positioning methods. The measuremen…

causal-discovery, classification, clustering, text

Computer Science

Image Segmentation

The instances were drawn randomly from a database of 7 outdoor images. The images were handsegmented to create a classification for every pixel. Ea…

classification, multivariate

Others

Extreme Classification Reposi…

The Extreme Classification Repository: Multi-label Datasets & Code Kush Bhatia Himanshu Jain Prateek Jain Manik Varma The objective in extreme mult…

benchmark, classification, evaluation, learning, machine, multilabel

Vision

Turkiye Student Evaluation

N/A

classification, clustering, multivariate

Others

Animals with attributes

A dataset for Attribute Based Classification. It consists of 30475 images of 50 animals classes with six pre-extracted feature representations for each im…

classification

Vision

BLOGGER

In this paper, we look for to recognize the causes of users tend to cyber space in Kohkiloye and Boyer Ahmad Province in Iran. Collecting information to f…

classification, multivariate

Computer Science

Connectionist Bench (Vowel Re…

The problem is specified by the accompanying data file, "vowel.data". This consists of a three dimensional array: voweldata [speaker, vowel, input]. The …

classification

Others

Middlebury MVS Temple

The object is a plaster reproduction of Temple of the Dioskouroi in Agrigento, Sicily. Click on thumbnail for a full-sized (640x480) image. Resolution of …

3d, 3d reconstruction, benchmark, multiview, sfm

Vision

Arcene

ARCENE was obtained by merging three mass-spectrometry datasets to obtain enough training and test data for a benchmark. The original features indicate th…

classification, multivariate

Life Sciences

CAMP-TUM: Multiple Human Pose…

We introduce the Shelf dataset for multiple human pose estimation from multiple views. In addition we annotate the body joints in the Campus dataset from …

3d, capture, estimation, human, motion, multiple, pose, view

Vision

CVC Partial Occlusion Virtual…

The CVC Partial Occlusion Virtual Pedestrian datasets (CVC-01 to CVC-06) cover a range of scenarios of occluded pedestrians generated in a virtual and rea…

classification, detection, occlusion, pedestrian, synthetic, tracking, urban

Vision

Breast Cancer Wisconsin (Diag…

Features are computed from a digitized image of a fine needle aspirate (FNA) of a breast mass. They describe characteristics of the cell nuclei present i…

classification, multivariate

Life Sciences

Google Street View Localizati…

The Google Street View dataset contains 62,058 high quality Google Street View images. The images cover the downtown and neighboring areas of Pittsburgh, …

address, google, gps, localization, manhattan, panorama, pittsburgh, retrieval, sphere, streetview, urban

Vision

Planning Relax

EEG record contains many regular oscillations, which are believed to reflect synchronized rhythmic activity in a group of neurons. Most activity related E…

classification, univariate

Computer Science

Mice Protein Expression

The data set consists of the expression levels of 77 proteins/protein modifications that produced detectable signals in the nuclear fraction of cortex. Th…

classification, clustering, multivariate

Life Sciences

OPPORTUNITY Activity Recognit…

The OPPORTUNITY Dataset for Human Activity Recognition from Wearable, Object, and Ambient Sensors is a dataset devised to benchmark human activity recogni…

classification, multivariate, time-series

Computer Science

Blood Transfusion Service Cen…

To demonstrate the RFMTC marketing model (a modified version of RFM), this study adopted the donor database of Blood Transfusion Service Center in Hsin-Ch…

classification, multivariate

Business

Legal Case Reports

This dataset contains Australian legal cases from the Federal Court of Australia (FCA). The cases were downloaded from AustLII ([Web Link]). We included a…

classification, text

Others

Bank Marketing

The data is related with direct marketing campaigns of a Portuguese banking institution. The marketing campaigns were based on phone calls. Often, more th…

classification, multivariate

Business

Online Handwritten Assamese C…

A dataset of online handwritten assamese characters by collecting samples from 45 writers is created. Each writer contributed 52 basic characters, 10 nume…

classification, multivariate, sequential

Computer Science

TVPR (Top View Person Re-iden…

The TVPR dataset includes 23 registration sessions. Each of the 23 folders contains the video of one registration session. Acquisitions have been performe…

clothing, depth, gender, identification, indoor, people, person, recognition, reidentification, top-view, video

Vision

SMS Spam Collection

This corpus has been collected from free or free for research sources at the Internet: -> A collection of 425 SMS spam messages was manually extracted fr…

classification, clustering, domain-theory, multivariate, text

Computer Science

CALTECH 256

The CALTECH 256 dataset by Li Fei-Fei contains 30607 images for 256 categories.

centered, classification, detection, image, object, scene

Vision

FMA: A Dataset For Music Anal…

* Audio track (encoded as mp3) of each of the 106,574 tracks. It is on average 10 millions samples per track.* Nine audio features (consisting of 518 attr…

classification, clustering, multivariate, time-series

Computer Science

Gesture Phase Segmentation

The dataset is composed by features extracted from 7 videos with people gesticulating, aiming at studying Gesture Phase Segmentation. Each video is repres…

classification, clustering, multivariate, sequential, time-series

Others

Paris Art Deco Facades

The Paris Art Deco Facades dataset consists of 79 / 80 images of rectified facades of the architectural style Art Deco, which has different sizes of windo…

architecture, city, facade, grammar, paris, procedural, recognition, segmentation, semantic, urban

Vision

Lane Level Localization on a …

The Lane Level Localization dataset was collected on a highway in San Francisco with the following properties: * Reasonable traffic * Multiple lane hig…

3d, autonomous, benchmark, car, driving, gps, localization, map, road, video

Vision

ISOLET

This data set was generated as follows. 150 subjects spoke the name of each letter of the alphabet twice. Hence, we have 52 training examples from each sp…

classification, multivariate

Computer Science

Stanford Dogs Dataset

The Stanford Dogs dataset contains images of 120 breeds of dogs from around the world. This dataset has been built using images and annotation from ImageN…

classification, detection, dogs, fine-grained categorization

Vision

Census Income

Extraction was done by Barry Becker from the 1994 Census database. A set of reasonably clean records was extracted using the following conditions: ((AAGE…

classification, multivariate

Social Sciences

Character Trajectories

The characters here were used for a PhD study on primitive extraction using HMM based models. The data consists of 2858 character samples, contained in th…

classification, clustering, time-series

Computer Science

Trains

Notes: - Additional "background" knowledge is supplied that provides a partial ordering on some of the attribute values. - We are providing this dataset…

classification, multivariate

Others

Urban scene recognition

Traffic Lights Recognition, Lara's public benchmarks.

recognition

Vision

Iris

This is perhaps the best known database to be found in the pattern recognition literature. Fisher's paper is a classic in the field and is referenced fre…

classification, multivariate

Life Sciences

Steel Plates Faults

Type of dependent variables (7 Types of Steel Plates Faults): 1.Pastry 2.Z_Scratch 3.K_Scatch 4.Stains 5.Dirtiness 6.Bumps 7.Other_Faults

classification, multivariate

Physical Systems

Hill-Valley

Each record represents 100 points on a two-dimensional graph. When plotted in order (from 1 through 100) as the Y co-ordinate, the points will create eith…

classification, sequential

Others

Heterogeneity Activity Recogn…

The Heterogeneity Dataset for Human Activity Recognition from Smartphone and Smartwatch sensors consists of two datasets devised to investigate sensor het…

classification, clustering, multivariate, time-series

Computer Science

Horse Colic

2 data files: -- horse-colic.data: 300 training instances -- horse-colic.test: 68 test instances Possible class attributes: 24 (whether lesi…

classification, multivariate

Life Sciences

Syskill and Webert Web Page R…

The HTML source of a web page is given. Users looked at each web page and inidated on a 3 point scale (hot medium cold) 50-100 pages per domain. However, …

classification, multivariate, text

Computer Science

Wilt

This data set contains some training and testing data from a remote sensing study by Johnson et al. (2013) that involved detecting diseased trees in Quick…

classification, multivariate

Life Sciences

Cervical cancer (Risk Factors)

The dataset was collected at 'Hospital Universitario de Caracas' in Caracas, Venezuela. The dataset comprises demographic information, habits, and histori…

classification, multivariate

Life Sciences

Dorothea

Drugs are typically small organic molecules that achieve their desired activity by binding to a target site on a receptor. The first step in the discovery…

classification, multivariate

Life Sciences

Ozone Level Detection

For a list of attributes, please refer to those two .names files. They use the following naming convention: All the attribute start with T means the tem…

classification, multivariate, sequential, time-series

Physical Systems

Daphnet Freezing of Gait

The Daphnet Freezing of Gait Dataset is a dataset devised to benchmark automatic methods to recognize gait freeze from wearable acceleration sensors plac…

classification, multivariate, time-series

Life Sciences

Gas sensor array exposed to t…

A chemical detection platform composed of 8 chemo-resistive gas sensors was exposed to turbulent gas mixtures generated naturally in a wind tunnel. The ac…

classification, multivariate, regression, time-series

Computer Science

Sentence Classification

Please see the README file that accompanies the data.

classification, text

Others

KUL Belgium Traffic Signs

BelgiumTS is a large dataset with 10000+ traffic sign annotations, thousands of physically distinct traffic signs. 4 video sequences recorded with 8 high …

belgium, calibration, camera, classification, road, sign, traffic, urban

Vision

KIMA99

The Kimia 99 has 9 classes each consisting of each 11 images. They are part of the Shape Indexing of Image Database (SIID) project, which also contains th…

binary, kimia, matching, shape retrieval

Vision

MSRC Kinect Gesture Dataset

The Microsoft Research Cambridge-12 Kinect gesture dataset consists of sequences of human movements, represented as body-part locations, and the associate…

action, gesture, human, kinect, recognition

Vision

Pedestrian Attribute Recognit…

Large-scale PEdesTrian Attribute (PETA) dataset, covering more than 60 attributes (e.g. gender, age range, hair style, casual/formal) on 19000 images.

recognition

Vision

Middlebury MVS Dino

The object is a plaster dinosaur (stegosaurus). Click on thumbnail for a full-sized (640x480) image. Resolution of ground truth model: 0.00025m (you may w…

3d, 3d reconstruction, benchmark, multiview, sfm

Vision

German Traffic Sign Recogniti…

The German Traffic Sign Recognition Benchmark is a dataset for multi-class detection problem in natural images and do cordially invite you to participate.…

detection, recognition, traffic, traffic sign, urban

Vision

Brodatz Album

The Brodatz dataset consists of 112 textures in grayscale images of various texture types. http://www.ee.oulu.fi/research/imag/texture/image_data/Brodat…

benchmark, classification, segmentation, synthetic, texture

Vision

Tic-Tac-Toe Endgame

This database encodes the complete set of possible board configurations at the end of tic-tac-toe games, where "x" is assumed to have played first. The t…

classification, multivariate

Games

ISTANBUL STOCK EXCHANGE

Data is collected from imkb.gov.tr and finance.yahoo.com. Data is organized with regard to working days in Istanbul Stock Exchange.

classification, multivariate, regression, time-series, univariate

Business

Dermatology

This database contains 34 attributes, 33 of which are linear valued and one of them is nominal. The differential diagnosis of erythemato-squamous diseas…

classification, multivariate

Life Sciences

gene expression cancer RNA-Seq

Samples (instances) are stored row-wise. Variables (attributes) of each sample are RNA-Seq gene expression levels measured by illumina HiSeq platform.

classification, clustering, multivariate

Life Sciences

Chess (King-Rook vs. King-Paw…

The dataset format is described below. Note: the format of this database was modified on 2/26/90 to conform with the format of all the other databases in…

classification, multivariate

Games

Activity Recognition system b…

This dataset represents a real-life benchmark in the area of Activity Recognition applications, as described in [1]. The classification tasks consist in…

classification, multivariate, sequential, time-series

Computer Science

Breast Cancer Wisconsin (Orig…

Samples arrive periodically as Dr. Wolberg reports his clinical cases. The database therefore reflects this chronological grouping of the data. This group…

classification, multivariate

Life Sciences

KTH Multiview Football

The KTH Multiview Football dataset contains 771 images of football players includes images taken from 3 views at 257 time instances 14 annotated body join…

camera, detection, game, multitarget, multiview, object, outdoor, pedestrian, pose, recognition, soccer, tracking

Vision

Eurasian Cities dataset

The Eurasian Cities dataset contains 103 images of outdoor urban scenes taken in Eurasian cities. It is annotated with horizontal and vertical vanishing p…

geometry, line, manhattan, outdoor, point, pose, reconstruction, urban, vanishing

Vision

TST TUG (Timed Up and Go)

The TUG (Timed Up and Go test) dataset consists of actions performed three times by 20 volunteers. The people involved in the test are aged between 22 and…

accelerometer, action, depth image processing - tug, human, kinect, recognition, time, video, wearable

Vision

M2CAI 2016 Challenge

These datasets were generated for the M2CAI challenges, a satellite event of MICCAI 2016 in Athens. Two datasets are available for two different challenge…

challenge, medicine, recognition, surgery, video, workflow

Vision

Yeast

Predicted Attribute: Localization site of protein. ( non-numeric ). The references below describe a predecessor to this dataset and its development. They…

classification, multivariate

Life Sciences

VidPairs

The VidPairs dataset contains 133 pairs of images, taken from 1080p HD (~2 megapixel) official movie trailers. Each pair consists of images of the same sc…

dense, description, flow, matching, optical, pair, patch, video

Vision

Abalone

Predicting the age of abalone from physical measurements. The age of abalone is determined by cutting the shell through the cone, staining it, and counti…

classification, multivariate

Life Sciences

ChokePoint Dataset

We collected a video dataset, termed ChokePoint, designed for experiments in person identification/verification under real-world surveillance conditions u…

clustering, detection, face, human, identification, multiview, pedestrian, real, recognition, sequence, surveillance, world

Vision

Firm-Teacher_Clave-Direction_…

The data consist of 16 binary inputs and one 'four-bit' one-hot classification output. The 16-bit inputs are binary-valued attack-point vectors. 1 indicat…

classification, multivariate

Others

HIV-1 protease cleavage

Past Usage: (a) Rgnvaldsson, You and Garwicz (2015) 'State of the art prediction of HIV-1 protease cleavage sites', Bioinformatics, vol 31 (8), p…

classification, multivariate

Life Sciences

Newspaper and magazine images…

This dataset was collected for training and validation of machine learning algorithm for classification regions of documents on text, picture and backgrou…

classification

Computer Science

Weight Lifting Exercises moni…

Velloso, E.; Bulling, A.; Gellersen, H.; Ugulino, W.; Fuks, H. Qualitative Activity Recognition of Weight Lifting Exercises. Proceedings of 4th Internatio…

classification, multivariate

Physical Systems

WordNet

WordNet is a large lexical database of English. Nouns, verbs, adjectives and adverbs are grouped into sets of cognitive synonyms (synsets), each expressin…

category, classification, hierarchy, imagenet, language

Vision

IMPART multi-modal/multi-view

The multi-modal/multi-view datasets are created in a cooperation between University of Surrey and Double Negative within the EU FP7 IMPART project. The …

3d, action, color, dynamic, emotion, face, human, indoor, lidar, model, multi-mode, multi-view, outdoor, rgbd, video

Vision

The KITTI Vision Benchmark Su…

We take advantage of our autonomous driving platform Annieway to develop novel challenging real-world computer vision benchmarks. Our tasks of interest ar…

depth, detection tracking, object detection, object tracking, odometry, optical flow, reconstruction, segmentation, semantic car depth, sfm, stereo

Vision

Zoo

A simple database containing 17 Boolean-valued attributes. The "type" attribute appears to be the class attribute. Here is a breakdown of which animals …

classification, multivariate

Life Sciences

Homepage

Description

Tags

Discussion

Related datasets