YACCLAB dataset

Vision Dataset

Homepage

http://imagelab.ing.unimore.it/imagelab/researchActivity.asp?idActivity=15

Description

The YACCLAB dataset includes both synthetic and real binary images and is suitable for a wide range of applications, ranging from document processing to survaillance, and features a significant variability in terms of resolution, image density and number of components. All images are provided in 1 bit per pixel PNG format, with 0 (black) being background and 1 (white) being foreground. Please include the following reference when citing the YACCLAB database: Grana, Costantino; Bolelli, Federico; Baraldi, Lorenzo; Vezzani, Roberto "YACCLAB - Yet Another Connected Components Labeling Benchmark" Proceedings of the 23rd International Conference on Pattern Recognition , Cancun, Mexico, 4-8 Dec 2016, 2016. Direct download: http://imagelab.ing.unimore.it/files/YACCLAB_dataset.zip Dataset page: http://imagelab.ing.unimore.it/imagelab/researchActivity.asp?idActivity=15

Discussion

Related datasets

UWO GCO Volume Segmentation

The Western GCO Segmentation problem instances are provided to compare effects of graph size, neighborhood size, length of s to t paths, regional arc cons…

abdomen, adhead, babyface, binary, bone, face, liver, medical, optimization, segmentation

Vision

SMS Spam Collection

This corpus has been collected from free or free for research sources at the Internet: -> A collection of 425 SMS spam messages was manually extracted fr…

classification, clustering, domain-theory, multivariate, text

Computer Science

Syskill and Webert Web Page R…

The HTML source of a web page is given. Users looked at each web page and inidated on a 3 point scale (hot medium cold) 50-100 pages per domain. However, …

classification, multivariate, text

Computer Science

Sentence Classification

Please see the README file that accompanies the data.

classification, text

Others

KIMA99

The Kimia 99 has 9 classes each consisting of each 11 images. They are part of the Shape Indexing of Image Database (SIID) project, which also contains th…

binary, kimia, matching, shape retrieval

Vision

UNIX User Data

This file contains 9 sets of sanitized user data drawn from the command histories of 8 UNIX computer users at Purdue over the course of up to 2 years (USE…

sequential, text

Computer Science

FIRE Fundus Image Registratio…

134 retinal image pairs and ground truth for registration.

medical

Vision

CALTECH 101 Category Patch Pa…

The CALTECH 101 Category Patch Pairs dataset measures invariance to intra-category variation. The dataset contains a training set and testing set of imag…

binary, feature description, feature matching, pair

Vision

SIID

The SIID silhouette dataset contains... and is from the Shape Indexing of Image Database (SIID). Download SIID silhouette dataset http://www.lems.brown…

binary, matching, shape retrieval

Vision

Mythological Creatures

The Mythological Creatures consists of articulated shapes (silhouettes) for partial similarity experiments and contains 15 shapes: 5 humans, 5 horses and …

animal, binary, matching, partial, shape retrieval

Vision

Sentiment Labelled Sentences

This dataset was created for the Paper 'From Group to Individual Labels using Deep Features', Kotzias et. al,. KDD 2015 Please cite the paper if you want …

classification, text

Others

ETHZ CVL Clust

MICCAI 2015 Challenge on Liver Ultrasound Tracking Munich, October 9, 2015 (Full Day) Outline Ultrasound (US) imaging is a widely used medical imaging…

benchmark, human, liver, medical, organ, real, therapy, tracking, ultrasound

Vision

Twenty Newsgroups

N/A

text

Others

YouTube Spam Collection

The table below lists the datasets, the YouTube video ID, the amount of samples in each class and the total number of samples per dataset. Dataset --- Yo…

classification, text

Computer Science

Burst Header Packet (BHP) flo…

For Further information about the variables see the file in the data folder.

classification, text

Computer Science

VIP Laparoscopic / Endoscopic…

Collection of endoscopic and laparoscopic (mono/stereo) videos and images

medical

Vision

Tools2D

The Tools 2D dataset from Bronstein, Bronstein, Bruckstein, and Kimmel [?] for partial similarity experiments and consists of 15 shapes: 5 humans, 5 horse…

binary, matching, partial, shape retrieval

Vision

Street View House Number (SVH…

SVHN is a real-world image dataset for developing machine learning and object recognition algorithms with minimal requirement on data preprocessing and fo…

classification, detection, number, real, recognition, streetside, streetview, text, urban, world

Vision

Twitter Data set for Arabic S…

--- By using a tweet crawler, we collect 2000 labelled tweets (1000 positive tweets and 1000 negative ones) on various topics such as: politics an…

classification, text

Social Sciences

Reuters Transcribed Subset

Data Characteristics: -------------------- This data was created by selecting 20 files each from the 10 largest classes in the Reuters-21578 collection …

classification, text

Business

KEGG Metabolic Reaction Netwo…

KEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. In …

classification, clustering, multivariate, regression, text, univariate

Life Sciences

Reuters-21578 Text Categoriza…

From the original readme file (please consult it for more information): ------------------------- The documents in the Reuters-21578 collection appeared o…

classification, text

Others

Reuter_50_50

The dataset is the subset of RCV1. These corpus has already been used in author identification experiments. In the top 50 authors (with respect to total s…

classification, clustering, domain-theory, multivariate, text

Computer Science

KIMA216

The Kimia 216 has 18 classes each consisting of 12 images. It contains shapes silhouettes for birds, bones, brick, camels, car, children, classic cards, e…

animal, binary, kimia, matching, shape retrieval

Vision

YouTube Multiview Video Games…

Please see the README for the details on the data organization, and so on.

classification, clustering, multivariate, text

Computer Science

Text and Vision (TVGraz) Data…

The Text and Vision (TVGraz) dataset is an annotated multi-modal dataset which currently contains 10 visual object categories, 4030 images and associated …

appearance, classification, evaluation, text

Vision

OpinRank Review Dataset

Car Reviews ------------ -Full reviews of cars for model-years 2007, 2008, and 2009 -There are about 140-250 cars for each model year -Extracted fields in…

text

Computer Science

Northix

Northix is designed to be a schema matching benchmark problem for data integration of two entity relationship databases. Northix is the resulting schema m…

classification, multivariate, text, univariate

Computer Science

TUD Shapes 1+2

This material is supplementary to Michael Stark, Bernt Schiele. How Good are Local Features for Classes of Geometric Objects. Eleventh IEEE Internatio…

binary, classification, object, shape, tool

Vision

KIMIA25

The Kimia 25 consists of 6 classes and 25 images. They are part of the Shape Indexing of Image Database (SIID) project, which also contains the SIID silho…

binary, kimia, matching, shape retrieval

Vision

microblogPCU

Our dataset is used by us to explore spammers in microblog and you can access our demo system at [Web Link]Please add :8080 after the domain name as port.…

causal-discovery, classification, multivariate, sequential, text, univariate

Computer Science

Dresses_Attribute_Sales

Style, Price, Rating, Size, Season, NeckLine, SleeveLength, waiseline, Material, FabricType, Decoration, Pattern, Type, Recommendation are Attributes in d…

classification, clustering, text

Computer Science

DBWorld e-mails

I collected 64 e-mails from DBWorld newsletter and I used them to train different algorithms in order to classify between 'announces of conferences' and '…

classification, text

Computer Science

Detail 2D Projection DataSet

Detail 2D Projection DataSet is a database of 2d projections of mechanical details with holes. The dataset consists of 13 shape categories where each cate…

binary, detail, holes, matching, shape retrieval

Vision

KEGG Metabolic Relation Netwo…

KEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. In …

classification, clustering, multivariate, regression, text, univariate

Life Sciences

MPEG-7 Core Experiment CE-Sha…

MPEG-7 Core Experiment CE-Shape-1 [?] is a popular database for shape matching evaluation consisting of 70 shape categories, where each category is repres…

binary, bullseye, matching, shape retrieval

Vision

NSF Research Award Abstracts …

The abstracts, one per file, were furnished by the NSF (National Science Foundation). A sample abstract is shown in the next section. The bag-of-word dat…

text

Others

Amazon Commerce reviews set

dataset are derived from the customers reviews in Amazon Commerce Website for authorship identification. Most previous studies conducted the identifi…

classification, domain-theory, multivariate, text

Physical Systems

KDC-4007 dataset Collection

The most important feature of this dataset is its simplicity to use and its being well-documented, which can be widely used in various studies of text ana…

classification, multivariate, regression, text

Computer Science

Mouse Embryo Tracking Database

DB Contains 100 examples with the uncompressed frames, up to the 10th frame after the appearance of the 8th cell; a text file with the trajectories of all…

medical

Vision

Eco-hotel

The CSV holds one cell per review, except the first row, which is the header. All the 401 reviews were collected between January and August of 2015.

text

Business

Tuberculosis image and patien…

Permanently growing database on lung tuberculosis patients. The data include radiological images (CT+XRay) plus social, clinical, and lab data as well as …

chest, ct, genome, medical, none, segmentation, tuberculosis, xray

Vision

Farm Ads

This data was collected from text ads found on twelve websites that deal with various farm animal related topics. Information from the ad creative and th…

classification, text

Business

Idiap/ETHZ Faces and Poses

Idiap/ETHZ Faces and Poses Dataset dataset by L. Jie, B. Caputo and V. Ferrari contains 1703 image-caption pairs. [author] Captions contain the names of s…

face detection, object pose, pedestrian, text

Vision

TTC-3600: Benchmark dataset f…

The dataset consists of a total of 3600 documents including 600 news/texts from six categories economy, culture-arts, health, politics, sports and techno…

classification, clustering, text

Computer Science

NIPS Conference Papers 1987-2…

The dataset is in the form of a 11463 x 5812 matrix of word counts, containing 11463 words and 5811 NIPS conference papers (the first column contains the …

clustering, text

Computer Science

Badges

Part of the problem in using an automated program to discover the unknown target function is to decide how to encode names such that the program can be us…

classification, text, univariate

Others

CNAE-9

This is a data set containing 1080 documents of free text business descriptions of Brazilian companies categorized into a subset of 9 categories cataloged…

classification, multivariate, text

Business

NYSK

Documents are first obtained via a Web search using AMIEI: an integrated platform for delivering enterprise intelligence, developed by AMI Software ([Web …

clustering, multivariate, sequential, text

Social Sciences

SML2010

The dataset could contain missing values. The data was sampled every minute, computing and uploading it smoothed with 15 minute means. The header of the d…

multivariate, regression, sequential, text, time-series

Computer Science

3D Road Network (North Jutlan…

This dataset was constructed by adding elevation information to a 2D road network in North Jutland, Denmark (covering a region of 185 x 135 km^2). Elevati…

clustering, regression, sequential, text

Computer Science

Opinosis Opinion Review

This dataset contains sentences extracted from user reviews on a given topic. Example topics are performance of Toyota Camry and sound quality of ipod nan…

text

Computer Science

Bag of Words

For each text collection, D is the number of documents, W is the number of words in the vocabulary, and N is the total number of words in the collection (…

clustering, text

Others

YouTube Comedy Slam Preferenc…

YouTube Comedy Slam ([Web Link]) is a video discovery experiment running on YouTube's version of labs (called TestTube) for a few months in 2011 and 2012.…

classification, text

Computer Science

Legal Case Reports

This dataset contains Australian legal cases from the Federal Court of Australia (FCA). The cases were downloaded from AustLII ([Web Link]). We included a…

classification, text

Others

Miskolc IIS Hybrid IPS

The measurements were created to ease the development, comparison and evaluation of fingerprinting based hybrid indoor positioning methods. The measuremen…

causal-discovery, classification, clustering, text

Computer Science

YACCLAB dataset

Homepage

Description

Tags

Discussion

Related datasets