An Annotated Dataset For Near-Duplicate Detection In Personal Photo Collections Managing photo collections involves a variety of image quality assessment tasks, e.g. the selection of the best photos. Detecting near-duplicates is a prerequisite for automating these tasks. We created the California-ND dataset to assist researchers in testing algorithms for the detection of near duplicate images. Contrary to other existing datasets in this domain, California-ND contains 701 photos taken directly from a real users personal photo collection. As a result, while including many challenging non-identical near-duplicate cases without the use of artificial image transformations. The original image sequence was maintained as much as possible. More importantly, in order to deal with the inevitable subjectivity and ambiguity that near-duplicate cases exhibit, the dataset is annotated by 10 different subjects, including the photographer himself. These annotations can be combined into a non-binary ground truth, representing the probability that a pair of images is considered a near-duplicate. http://vintage.winklerbros.net/californiaND.zip A. Jinda-Apiraksa, V. Vonikakis, S. Winkler. California-ND: An annotated dataset for near-duplicate detection in personal photo collections. Proc. 5th International Workshop on Quality of Multimedia Experience (QoMEX), Klagenfurt, Austria, July 3-5, 2013.California-ND is a dataset for non-identical near-duplicates which contains 701 travel photos annotated by 10 subjects, including the photographer. Managing photo collections involves a variety of image quality assessment tasks, e. the selection of the "best" photos. The California-ND dataset assists researchers in testing algorithms for the detection of near duplicate images. California-ND contains 701 photos taken directly from a real user's personal photo collection.
Some datasets and evaluation tools are provided on this page for four different computer vision and computer graphics problems. Population counting L...
urban, surface, reconstruction, pointcloud, object, road, pedestrian, network, line, 3d, crowd, counting, detection, groundtruthThe mirror symmetry database contains 176 single-symmetry and 63 multyple-symmetry images (.png files) with accompanying ground-truth annotations (.mat ...
symmetry, detection, groundtruth, mirrorThe FlickrLogos-32 dataset contains photos showing brand logos and is meant for the evaluation of multi-class logo recognition as well as logo retrieval...
classification brand boundingbox, retrieval, object recognition, machine learning, logo, detection, image, flickrThe Ford Car dataset is joint effort of Pandey et al. (for collecting images, Lidar points, calibration etc.) and us (for annotation of 2D and 3D object...
lidar, detection, groundtruth, 3d, car, sfmLASIESTA is composed by many real indoor and outdoor sequences organized in different categories, each of one covering a specific challenge in moving ob...
motion, subtraction, dataset, background, object, stationary, foreground, camera, challenge, detection, groundtruthCollected in a clothing store. Captured with Kinect (640*480, about 30fps)
tracking, detectionThe FaceScrub dataset comprises a total of 107818 unconstrained face images of 530 celebrities crawled from the Internet, with about 200 images per pers...
face, celebrity, detection, people, recognition, humanWe introduce a labeled dataset of categorized images for evaluating sketch based image retrieval. Using Flickr, we downloaded about 3000 images for each...
saliency, internet, shape, sketch, visual, attention, group, retrieval, salient object detectionThe Oxford RobotCar Dataset contains over 100 repetitions of a consistent route through Oxford, UK, captured over a period of over a year. The dataset c...
driving, street, urban, time, recognition, autonomous, video, segmentation, robot, classification, detection, car, yearThe ImageNET dataset is the latest dataset by Li Fei-Fei containing various dataset ranging from 1000 to 10000 categories.
image classification, object segmentation, retrievalThe QMUL Junction dataset is a busy traffic scenario for research on activity analysis and behavior understanding. Video length: 1 hour (90000 frame...
video, motion, pedestrian, crowd, counting, tracking, detection, behaviorThe ICG Lab 6 (Multi-Camera Multi-Object Tracking) dataset contains 6 indoor people tracking scenarios recorded at our laboratory using 4 static Axis P1...
evaluation, graz, object, laboratory, pedestrian, segmentation, multiview, tracking, camera, detection, calibrationBackground Models Challenge (BMC) is a complete dataset and competition for the comparison of background subtraction algorithms. The main topics concer...
motion, background, video, modeling, segmentation, change, surveillance, detectionThe Video Segmentation Benchmark (VSB100) provides ground truth annotations for the Berkeley Video Dataset, which consists of 100 HD quality videos divi...
video, object, segmentation, motion, pedestrian, benchmark, tracking, groundtruthThe Daimler Mono Pedestrian Detection Benchmark dataset contains a large training and test set. The training set contains 15.560 pedestrian samples (ima...
object, mono, urban, pedestrian, outdoor, scale, detectionPhos is a color image database of 15 scenes captured under different illumination conditions. More particularly, every scene of the database contains 1...
detectionThe HandNet dataset contains depth images of 10 participants hands non-rigidly deforming infront of a RealSense RGB-D camera. This dataset includes 2...
rgbd, hand, articulation, video, segmentation, classification, pose, fingertip, detectionThe TUD Crossing dataset from Micha Andriluka, Stefan Roth and Bernt Schiele consists of 201 images with 1008 highly overlapping pedestrians with signif...
urban, sideview, overlap, segmentation, pedestrian, tracking, multitarget, detectionThe Caltech Buildings dataset consists of images taken for 50 buildings around the Caltech campus. Five different images were taken for each building fr...
building, caltech, urban, retrieval, taxonomy, hierarchyThis UIUC Cars dataset by Shivani Agarwal, Aatif Awan and Dan Roth contains images of side views of cars for use in evaluating object detection algorith...
urban, sideview, detection, car, recognition, scaleThe dataset captures 25 people preparing 2 mixed salads each and contains over 4h of annotated accelerometer and RGB-D video data. Annotated activities ...
video, activity, classification, tracking, recognition, detection, actionThe UCF Person and Car VideoSeg dataset consists of six videos with groundtruth for video object segmentation. Surfing, jumping, skiing, sliding, big ...
video, object, segmentation, motion, model, camera, groundtruthThe UrbanStreet dataset used in the paper can be downloaded here [188M] . It contains 18 stereo sequences of pedestrians taken from a stereo rig mounted...
urban, human, recognition, video, pedestrian, segmentation, tracking, multitarget, detectionScene Background Initialization (SBI) dataset The SBI dataset has been assembled in order to evaluate and compare the results of background initializa...
change, detection, benchmark, background, foreground, initializationThe Traffic Video dataset consists of X video of an overhead camera showing a street crossing with multiple traffic scenarios. The dataset can be down...
video, urban, traffic, road, overhead, tracking, view, detectionSince the publicly available face image datasets are often of small to medium size, rarely exceeding tens of thousands of images, and often without age ...
face, age, wikipedia, imdb, recognition, detection, biometryThe Webcam Interestingness dataset consists of 20 different webcam streams, with 159 images each. It is annotated with interestingness ground truth, acq...
video, interest, retrieval, classification, weather, ranking, webcamThe High Definition Analytics (HDA) dataset is a multi-camera High-Resolution image sequence dataset for research on High-Definition surveillance: Pedes...
high-definition, benchmark, human, lisbon, indoor, video, re-identification, pedestrian, network, multiview, tracking, surveillance, camera, detectionThe Stanford 40 Actions dataset contains images of humans performing 40 actions. In each image, we provide a bounding box of the person who is performin...
recognition, human, detection, action, boundingboxThis dataset contains 12,995 face images which are annotated with (1) five facial landmarks, (2) attributes of gender, smiling, wearing glasses, and hea...
face, landmark detection, deep learning, detection, attribute, cnnThe Google Street View dataset contains 62,058 high quality Google Street View images. The images cover the downtown and neighboring areas of Pittsburgh...
pittsburgh, urban, manhattan, sphere, address, panorama, google, streetview, gps, retrieval, localizationAt Udacity, we believe in democratizing education. How can we provide opportunity to everyone on the planet? We also believe in teaching really amazing ...
driving, street, urban, time, recognition, autonomous, video, segmentation, robot, classification, detection, car, syntheticThe Leeds Cows dataset by Derek Magee consists of 14 different video sequences showing a total of 18 cows walking from right to left in front of differe...
video, segmentation, detection, cow, animal, backgroundThe TRaffic ANd COngestionS (TRANCOS) dataset, a novel benchmark for (extremely overlapping) vehicle counting in traffic congestion situations. It consi...
urban, highway, spain, object, traffic, transportation, vehicle, detection, carThe Annotated Facial Landmarks in the Wild (AFLW) consists of a large-scale collection of annotated face images gathered from the web, exhibiting a larg...
face, annotation, detection, age, landmark, poseThe German Traffic Sign Recognition Benchmark is a dataset for multi-class detection problem in natural images and do cordially invite you to participat...
urban, traffic, recognition, detection, traffic signThe CALTECH 256 dataset by Li Fei-Fei contains 30607 images for 256 categories.
object, detection, image, centered, classification, sceneThe Wide (multiple) Baseline Dataset. 31 image pairs, simultaneously combining several nuisance factors: geometry, illumination, IR-visible, etc. WxBS...
description, night, viewpoint, matching, feature, detection, day, irContains 6 object categories similar to object categories in Pascal VOC that are suitable for studying the abnormalities stemming from objects.
detection1521 images with human faces, recorded under natural conditions, i.e. varying illumination and complex background. The eye positions have been set manua...
detectionWIDER FACE dataset is a large-scale face detection benchmark dataset with 32,203 images and 393,703 face annotations, which have high degree of variabil...
face, scale, detection, pose, occlusionThe PETS 2009 dataset contains 3 parts showing multi-view sequences containing pedestrians walking in an outdoor environment. The parts are used for per...
overlap, human, frontview, occlusion multitarget, outdoor, pedestrian, tracking, detectionThe contour patches dataset is a large dataset of images patch matches used for contour detection. References: C. L. Zitnick and D. Parikh The Role...
lowlevel, match, edge, image, contour, segmentation, patch, detectionThe Inria Aerial Image Labeling addresses a core topic in remote sensing: the automatic pixelwise labeling of aerial imagery (link to paper). Dataset ...
house, urban, aerial, building, segmentation, footprint, groundtruth, city, semanticThe TU Berlin Multi-Object and Multi-Camera Tracking Dataset (MOCAT) is a synthetic dataset to train and test tracking and detection systems in a virtua...
evaluation, multi-view, pedestrian, animal, tracking, multi-class, vehicle, detection, syntheticThe Caltech Game Covers dataset consists of CD/DVD covers of video games. The set was downloaded from freecovers.net during the summer of 2008. The set ...
caltech, retrieval, game, cover, classification, hierarchy, taxonomyThe Swedish Traffic Sign Recognition provides Matlab code for parsing the annotation files and displaying the results. Part0 for each set contains the a...
urban, traffic, detection, city, sign, recognitionCOCO-Stuff augments the COCO dataset with pixel-level stuff annotations for 10,000 images. These annotations can be used for scene understanding tasks l...
annotation, benchmark, coco, segmentation, things, captioning, stuff, groundtruth, semanticSVHN is a real-world image dataset for developing machine learning and object recognition algorithms with minimal requirement on data preprocessing and ...
urban, real, recognition, text, streetside, world, streetview, classification, detection, numberThe Aspect Layout dataset is designed to allow evaluation of object detection for aspect ratios in perspective images. Author text: In this project ...
object, detection, aspect, perspective, ratio, layoutThe Quad 6K dataset is a Structure-from-Motion dataset taken at Arts Quad at Cornell University campus and consists of 6514 images with ground truth pos...
urban, 3d reconstruction, groundtruth, sfm, landmark, 3d gpsThe San Francisco Landmark Dataset for Mobile Landmark Recognition is a set of images and query images for localization. We present the San Francisco ...
urban, mobile, sanfrancisco, gps, retrieval, localization, landmark, city, calibrationThis dataset contains 7 challenging volleyball activity classes annotated in 6 videos from professionals in the Austrian Volley League (season 2011/12)....
video, sport, analysis, activity recognition, volleyball, detection, actionContains drawing pages from US patents with manually labeled figure and part labels.
detectionTh EPFL Multi-View Car dataset contains 20 sequences of cars as they rotate by 360 degrees. There is one image approximately every 3-4 degrees. Using th...
detection, estimation, car, pose, multiview, rotationThe 1DSfM Landmarks is a collection of community-based image reconstruction by Kyle Wilson and is comprised of 14 datasets with comparison to bundler gr...
urban, 3d, benchmark, city, reconstruction, landmark, groundtruthThe PASCAL VOC is augmented with segmentation annotation for semantic parts of objects. For example, for the person category, we provide segmentation ma...
part, human, recognition, object, pedestrian, segmentation, pascal, detection, semanticLabelMe is a web-based image annotation tool that allows researchers to label images and share the annotations with the rest of the community. If you us...
detectionThe YouTube-Objects dataset is composed of videos collected from YouTube by querying for the names of 10 object classes. It contains between 9 and 24 vi...
video, object, flow, segmentation, detection, opticalYahoo Flickr Creative Commons 100M (YFCC100M) dataset contains a list of photos and videos. This list is compiled from data available on Yahoo! Flickr. ...
internet, reconstruction, recognition, image, community, social, 3d, clustering, detection, flickr, landmarkThe KTH Multiview Football dataset contains 771 images of football players includes images taken from 3 views at 257 time instances 14 annotated body jo...
recognition, soccer, outdoor, object, pedestrian, game, pose, multiview, tracking, camera, multitarget, detectionA New Color Image Database for Benchmarking of Face Detection Techniques and Human Skin Segmentation Techniques. A new color face image database for ...
face, segmentation, skin, detection, benchmarkingGlobal Symmetry Ground-truth for AVA dataset Release Date: 2016 For detailed information, please refer to: Elawady, Mohamed, Ccile Barat, Christoph...
bilateral, aesthetic, global, symmetry, reflection, detection, mirrorWe collected a video dataset, termed ChokePoint, designed for experiments in person identification/verification under real-world surveillance conditions...
face, real, human, recognition, world, pedestrian, identification, clustering, multiview, surveillance, detection, sequence10000 images of natural scenes, with 37 different logos, and 2695 logos instances, annotated with a bounding box.
detectionThe SegTrack dataset consists of six videos (five are used) with ground truth pixelwise segmentation (6th penguin is not usable). The dataset is used fo...
motion, video, object, proposal, flow, segmentation, stationary, model, camera, optical, groundtruthThe CERTH image blur dataset consists of 2450 digital images, 1850 out of which are photographs captured by various camera models in different shooting ...
motion, quality, detection, image, defocus, blurThe crowd datasets are collected from a variety of sources, such as UCF and data-driven crowd datasets. The sequences are diverse, representing dense cr...
video, pedestrian, scene, crowd, human, understanding, anomaly, detectionThe Microsoft COCO (mscoco) is an image recognition and segmentation dataset which contains more 300k images for more than 70 categories. Other featur...
object, segmentation, benchmark, semantic, context, recognition, detectionThe MSR Action datasets is a collection of various 3D datasets for action recognition. See details http://research.microsoft.com/en-us/um/people/zliu...
video, detection, 3d, action, reconstruction, recognitionThe Mall dataset was collected from a publicly accessible webcam for crowd counting and profiling research. Ground truth: Over 60,000 pedestrians wer...
video, pedestrian, crowd, counting, tracking, detection, indoor, webcamThe Where Who Why (WWW) dataset provides 10,000 videos with over 8 million frames from 8,257 diverse scenes, therefore offering a superior comprehensive...
recognition, video, flow, pedestrian, crowd, surveillance, optical, detection10000 images of natural scenes grabbed on Flickr, with 2695 logos instances cut and pasted from the BelgaLogos dataset.
detectionMany different labeled video datasets have been collected over the past few years, but it is hard to compare them at a glance. So we have created a hand...
video, object, benchmark, classification, recognition, detection, action15 wide baseline stereo image pairs with large viewpoint change, provided ground truth homographies. Image size (~1000x700 pixels, RGB) D. Mishkin a...
description, wide baseline stereo, detection, viewpoint, matching, featureThe Freiburg-Berkeley Motion Segmentation Dataset (FBMS-59) is an extension of the BMS dataset with 33 additional video sequences. A total of 720 frames...
motion, benchmark, video, object, pedestrian, segmentation, tracking, groundtruthThe CVC Partial Occlusion Virtual Pedestrian datasets (CVC-01 to CVC-06) cover a range of scenarios of occluded pedestrians generated in a virtual and r...
urban, pedestrian, classification, synthetic, occlusion, tracking, detectionClassification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
detection, classificationThe PETS 2016 IPATCH dataset contains a set of fourteen multi camera recordings (visible, themal) collected off the coast of Brest, France, in collabora...
visible, thermal, multimodal, vessel, maritime, boat, gps, tracking, detection, radarThe Caltech Lanes dataset includes four clips taken around streets in Pasadena, CA at different times of day. The archive below includes 1225 individu...
caltech, urban, road, pasadena, detection, laneMultispectral Imaging (MSI) datasets were acquired using IRIS II which is a lightweight portable system comprising of a high resolution camera, a novel ...
illumination, wavelength, registration, alignment, matching, groundtruth, multi-spectralThe Video Summarization (SumMe) dataset consists of 25 videos, each annotated with at least 15 human summaries (390 in total). The data consists of vide...
video, benchmark, summary, event, human, groundtruth, action30000+ frames with vehicle rear annotation and classification (car and trucks) on motorway/highway sequences. Annotation semi-automatically generated us...
detectionThe Multi-FoV synthetic datasets are two synthetic scenes (vehicle moving in a city, and flying robot hovering in a confined room). For each scene, thre...
synthetic, visual, odometry, fov, blender, camera, groundtruthThe Compact Descriptors for Visual Search Patches Dataset (CDVS) is a dataset comprised of pairwise image patches. MPEG is a standard titled Compact De...
descriptor, mpeg, patch, retrieval, matching, featureThe Stanford Dogs dataset contains images of 120 breeds of dogs from around the world. This dataset has been built using images and annotation from Imag...
fine-grained categorization, dogs, detection, classificationChairGest is an open challenge / benchmark. The task consists in spotting and recognizing gestures from multiple synchronized sensors: 1 Kinect and 4 X...
gesture, detection, benchmark, kinect, recognition, humanThis dataset provides a collection of web images and 3D models for research on landmark recognition (especially for methods based on 3D models). We hope...
codebook, reconstruction, matching, recognition, retrieval, 3d, classification, feature, flickr, landmarkThe PIROPO database (People in Indoor ROoms with Perspective and Omnidirectional cameras) comprises multiple sequences recorded in two different indoor ...
perspective, human, indoor, room, surveillance, detection, fisheye, omnidirectional, peopleWe present a new large-scale dataset that contains a diverse set of stereo video sequences recorded in street scenes from 50 different cities, with high...
urban, stereo, cities, person, video, weakly, segmentation, pedestrian, detection, car, semanticPenn-Fudan Pedestrian Detection and Segmentation
segmentation, motion, background, pedestrian, detectionThe Kendall Square webcam dataset consists of two streams for one sunny day and one cloudy day of a city square. It is used for tracking and analyzing c...
color, change, appearance, weather, detection, webcam, skyt is composed of food intake movements, recorded with Kinect V1 (320240 depth frame resolution), simulated by 35 volunteers for a total of 48 tests. The...
kinect, age, intake, pointcloud, human, tracking, monitoring, groundtruth, food, behaviorThe city planar and non-planar datset consists of urban scenes accompanied by text files describing the plane/non-plane locations. Training Set (Univ...
building, urban, detection, 3d, estimation, planeThe ICG Multi-Camera and Virtual PTZ dataset contains the video streams and calibrations of several static Axis P1347 cameras and one panoramic video fr...
graz, outdoor, video, object, panorama, pedestrian, network, crowd, multiview, tracking, camera, multitarget, detection, calibrationThe Extreme Zoom Dataset. EZD is a 6 image sets with incleasing zoom factor from general scene view to focusing on single detail. MODS: Fast and Robus...
description, detection, zoom, viewpoint, matching, featureThe dataset contains 15 documentary films that are downloaded from YouTube, whose durations vary from 9 minutes to as long as 50 minutes, and the total ...
video, object, detectionInstance recognition from depth data. Contains various challenges of Pose, Clutter, Occlusion and similar looking objects (Bonde, U., Badrinarayanan, V....
detection, instance, depth, poseWe share our omnidirectional and panoramic image dataset (with annotations) to be used for human and car detection. Please reach through: http://cvrg.i...
panorama, detection, car, omnidirection, recognition, humanThe Longterm Pedestrian dataset consists of images from a stationary camera running 24 hours for 7 days at about 1 fps. It used for adaptive detection ...
coffee, graz, background, indoor, illumination, change, pedestrian, robust, multitarget, detection15,560 pedestrian and non-pedestrian samples (image cut-outs) and 6744 additional full images not containing pedestrians for bootstrapping. The test set...
detectionCalifornia-ND contains 701 photos taken directly from a real user's personal photo collection, including many challenging non-identical near-duplicate c...
detectionA 66 stereo pairs dataset with their subpixel ground truths. The construction and improvement of algorithms for subpixel stereovision requires very pr...
stereo, depth, pointcloud, noise, stereovision, 3d, groundtruth, subpixelThe ICG Multi-Camera datasets consist of Easy Data Set (just one person) Medium Data Set (3-5 persons, used for the experiments) Hard Data Set (cro...
graz, indoor, video, object, pedestrian, multiview, tracking, camera, multitarget, detection, calibrationThe CMP map2photo dataset consists of 6 pairs, where one image is satellite photo and second image is a map of the same area. The task is to match thes...
sensing, baseline, matching, description, map, feature, remote, detection, wide