UNIX User Data

Computer Science Dataset

Homepage

http://archive.ics.uci.edu/ml/datasets/UNIX+User+Data

Description

This file contains 9 sets of sanitized user data drawn from the command histories of 8 UNIX computer users at Purdue over the course of up to 2 years (USER0 and USER1 were generated by the same person, working on different platforms and different projects). The data is drawn from tcsh(1) history files and has been parsed and sanitized to remove filenames, user names, directory structures, web addresses, host names, and other possibly identifying items. Command names, flags, and shell metacharacters have been preserved. Additionally, **SOF** and **EOF** tokens have been inserted at the start and end of shell sessions, respectively. Sessions are concatenated by date order and tokens appear in the order issued within the shell session, but no timestamps are included in this data. For example, the two sessions: # Start session 1 cd ~/private/docs ls -laF | more cat foo.txt bar.txt zorch.txt > somewhere exit # End session 1 # Start session 2 cd ~/games/ xquake & fg vi scores.txt mailx john_doe '@' somewhere.com exit # End session 2 would be represented by the token stream **SOF** cd <1> # one "file name" argument ls -laF | more cat <3> # three "file" arguments > <1> exit **EOF** **SOF** cd <1> xquake & fg vi <1> mailx <1> exit **EOF**

Discussion

Related datasets

SML2010

The dataset could contain missing values. The data was sampled every minute, computing and uploading it smoothed with 15 minute means. The header of the d…

multivariate, regression, sequential, text, time-series

Computer Science

3D Road Network (North Jutlan…

This dataset was constructed by adding elevation information to a 2D road network in North Jutland, Denmark (covering a region of 185 x 135 km^2). Elevati…

clustering, regression, sequential, text

Computer Science

microblogPCU

Our dataset is used by us to explore spammers in microblog and you can access our demo system at [Web Link]Please add :8080 after the domain name as port.…

causal-discovery, classification, multivariate, sequential, text, univariate

Computer Science

NYSK

Documents are first obtained via a Web search using AMIEI: an integrated platform for delivering enterprise intelligence, developed by AMI Software ([Web …

clustering, multivariate, sequential, text

Social Sciences

Gesture Phase Segmentation

The dataset is composed by features extracted from 7 videos with people gesticulating, aiming at studying Gesture Phase Segmentation. Each video is repres…

classification, clustering, multivariate, sequential, time-series

Others

Hill-Valley

Each record represents 100 points on a two-dimensional graph. When plotted in order (from 1 through 100) as the Y co-ordinate, the points will create eith…

classification, sequential

Others

Syskill and Webert Web Page R…

The HTML source of a web page is given. Users looked at each web page and inidated on a 3 point scale (hot medium cold) 50-100 pages per domain. However, …

classification, multivariate, text

Computer Science

Ozone Level Detection

For a list of attributes, please refer to those two .names files. They use the following naming convention: All the attribute start with T means the tem…

classification, multivariate, sequential, time-series

Physical Systems

Sentence Classification

Please see the README file that accompanies the data.

classification, text

Others

Activity Recognition system b…

This dataset represents a real-life benchmark in the area of Activity Recognition applications, as described in [1]. The classification tasks consist in…

classification, multivariate, sequential, time-series

Computer Science

Molecular Biology (Protein Se…

This is a data set used by Ning Qian and Terry Sejnowski in their study using a neural net to predict the secondary structure of certain globular proteins…

classification, sequential

Life Sciences

Activities of Daily Living (A…

This dataset comprises information regarding the ADLs performed by two users on a daily basis in their own homes. This dataset is composed by two instanc…

classification, clustering, multivariate, sequential, time-series

Computer Science

Cargo 2000 Freight Tracking a…

A description of the underlying Cargo 2000 standard and the processes reflected in the data set can be found at [Web Link].

classification, multivariate, regression, sequential

Business

Sentiment Labelled Sentences

This dataset was created for the Paper 'From Group to Individual Labels using Deep Features', Kotzias et. al,. KDD 2015 Please cite the paper if you want …

classification, text

Others

Twenty Newsgroups

N/A

text

Others

YACCLAB dataset

The YACCLAB dataset includes both synthetic and real binary images and is suitable for a wide range of applications, ranging from document processing to s…

binary, fingerprints, labeling, medical, natural, randomnoise, text, videosurveillance

Vision

Grammatical Facial Expressions

The automated analysis of facial expressions has been widely used in different research areas, such as biometrics or emotional analysis. Special import…

classification, clustering, multivariate, sequential

Computer Science

YouTube Spam Collection

The table below lists the datasets, the YouTube video ID, the amount of samples in each class and the total number of samples per dataset. Dataset --- Yo…

classification, text

Computer Science

Burst Header Packet (BHP) flo…

For Further information about the variables see the file in the data folder.

classification, text

Computer Science

Street View House Number (SVH…

SVHN is a real-world image dataset for developing machine learning and object recognition algorithms with minimal requirement on data preprocessing and fo…

classification, detection, number, real, recognition, streetside, streetview, text, urban, world

Vision

Twitter Data set for Arabic S…

--- By using a tweet crawler, we collect 2000 labelled tweets (1000 positive tweets and 1000 negative ones) on various topics such as: politics an…

classification, text

Social Sciences

Wearable Computing: Classific…

IMPORTANT: we have lower performance on 'leave-one-subject-out' tests. The performance baseline index we established is for 10-fold cross-validation tests…

classification, sequential

Computer Science

QtyT40I10D100K

This data set is generated from the original T40I10D100K data set, to mine fuzzy sequential patterns over quantitative streams. While the original T40I10D…

sequential

Others

Reuters Transcribed Subset

Data Characteristics: -------------------- This data was created by selecting 20 files each from the 10 largest classes in the Reuters-21578 collection …

classification, text

Business

KEGG Metabolic Reaction Netwo…

KEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. In …

classification, clustering, multivariate, regression, text, univariate

Life Sciences

Reuters-21578 Text Categoriza…

From the original readme file (please consult it for more information): ------------------------- The documents in the Reuters-21578 collection appeared o…

classification, text

Others

Molecular Biology (Splice-jun…

Problem Description: Splice junctions are points on a DNA sequence at which `superfluous' DNA is removed during the process of protein creation in highe…

classification, domain-theory, sequential

Life Sciences

Reuter_50_50

The dataset is the subset of RCV1. These corpus has already been used in author identification experiments. In the top 50 authors (with respect to total s…

classification, clustering, domain-theory, multivariate, text

Computer Science

User Identification From Walk…

The dataset collects data from an Android smartphone positioned in the chest pocket. Accelerometer Data are collected from 22 participants walking in the …

classification, clustering, sequential, time-series, univariate

Others

Online Retail

This is a transnational data set which contains all the transactions occurring between 01/12/2010 and 09/12/2011 for a UK-based and registered non-store o…

classification, clustering, multivariate, sequential, time-series

Business

Molecular Biology (Promoter G…

This dataset has been developed to help evaluate a "hybrid" learning algorithm ("KBANN") that uses examples to inductively refine preexisting knowledge. …

classification, domain-theory, sequential

Life Sciences

YouTube Multiview Video Games…

Please see the README for the details on the data organization, and so on.

classification, clustering, multivariate, text

Computer Science

Text and Vision (TVGraz) Data…

The Text and Vision (TVGraz) dataset is an annotated multi-modal dataset which currently contains 10 visual object categories, 4030 images and associated …

appearance, classification, evaluation, text

Vision

EEG Eye State

All data is from one continuous EEG measurement with the Emotiv EEG Neuroheadset. The duration of the measurement was 117 seconds. The eye state was detec…

classification, multivariate, sequential, time-series

Life Sciences

OpinRank Review Dataset

Car Reviews ------------ -Full reviews of cars for model-years 2007, 2008, and 2009 -There are about 140-250 cars for each model year -Extracted fields in…

text

Computer Science

Northix

Northix is designed to be a schema matching benchmark problem for data integration of two entity relationship databases. Northix is the resulting schema m…

classification, multivariate, text, univariate

Computer Science

Educational Process Mining (E…

The experiments have been carried out with a group of 115 students of first-year, undergraduate Engineering major of the University of Genoa. We carried…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

Predict keywords activities i…

See files and/or [Web Link]

multivariate, sequential, time-series

Computer Science

Dresses_Attribute_Sales

Style, Price, Rating, Size, Season, NeckLine, SleeveLength, waiseline, Material, FabricType, Decoration, Pattern, Type, Recommendation are Attributes in d…

classification, clustering, text

Computer Science

DBWorld e-mails

I collected 64 e-mails from DBWorld newsletter and I used them to train different algorithms in order to classify between 'announces of conferences' and '…

classification, text

Computer Science

KEGG Metabolic Relation Netwo…

KEGG Metabolic pathways can be realized into network. Two kinds of network / graph can be formed. These include Reaction Network and Relation Network. In …

classification, clustering, multivariate, regression, text, univariate

Life Sciences

UJI Pen Characters

We create a character database by collecting samples from 11 writers. Each writer contributed with letters (lower and uppercase), digits, and other chara…

classification, multivariate, sequential

Computer Science

NSF Research Award Abstracts …

The abstracts, one per file, were furnished by the NSF (National Science Foundation). A sample abstract is shown in the next section. The bag-of-word dat…

text

Others

Wall-Following Robot Navigati…

The provided files comprise three different data sets. The first one contains the raw values of the measurements of all 24 ultrasound sensors and the cor…

classification, multivariate, sequential

Computer Science

Hybrid Indoor Positioning Dat…

The measurements were created to ease the development, comparison and evaluation of fingerprinting based hybrid indoor positioning methods. The measuremen…

classification, multivariate, sequential, time-series

Computer Science

Amazon Commerce reviews set

dataset are derived from the customers reviews in Amazon Commerce Website for authorship identification. Most previous studies conducted the identifi…

classification, domain-theory, multivariate, text

Physical Systems

UJIIndoorLoc-Mag

Indoor localization is a key topic for mobile computing. However, it is still very difficult for the mobile sensing community to compare state-of-art Indo…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

Geo-Magnetic field and WLAN d…

Indoor localisation is a key topic for the Ambient Intelligence (AmI) research community. In this scenarios, recent advancements in wearable technologie…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

MSNBC.com Anonymous Web Data

The data comes from Internet Information Server (IIS) logs for msnbc.com and news-related portions of msn.com for the entire day of September, 28, 1999 (P…

sequential

Computer Science

KDC-4007 dataset Collection

The most important feature of this dataset is its simplicity to use and its being well-documented, which can be widely used in various studies of text ana…

classification, multivariate, regression, text

Computer Science

Bach Choral Harmony

Pitch classes information has been extracted from MIDI sources downloaded from (JSB Chorales)[[Web Link]]. Meter information has been computed through the…

classification, sequential

Others

Eco-hotel

The CSV holds one cell per review, except the first row, which is the header. All the 401 reviews were collected between January and August of 2015.

text

Business

Libras Movement

The dataset (movement_libras) contains 15 classes of 24 instances each, where each class references to a hand movement type in LIBRAS. In the video pre-p…

classification, clustering, multivariate, sequential

Others

Farm Ads

This data was collected from text ads found on twelve websites that deal with various farm animal related topics. Information from the ad creative and th…

classification, text

Business

Activity Recognition from Sin…

--- The dataset collects data from a wearable accelerometer mounted on the chest --- Sampling frequency of the accelerometer: 52 Hz --- Accelerom…

classification, clustering, sequential, time-series, univariate

Others

Idiap/ETHZ Faces and Poses

Idiap/ETHZ Faces and Poses Dataset dataset by L. Jie, B. Caputo and V. Ferrari contains 1703 image-caption pairs. [author] Captions contain the names of s…

face detection, object pose, pedestrian, text

Vision

TTC-3600: Benchmark dataset f…

The dataset consists of a total of 3600 documents including 600 news/texts from six categories economy, culture-arts, health, politics, sports and techno…

classification, clustering, text

Computer Science

Taxi Service Trajectory - Pre…

For complete information see the official challenge page: [Web Link]

causal-discovery, clustering, domain-theory, multivariate, sequential, time-series

Computer Science

NIPS Conference Papers 1987-2…

The dataset is in the form of a 11463 x 5812 matrix of word counts, containing 11463 words and 5811 NIPS conference papers (the first column contains the …

clustering, text

Computer Science

Localization Data for Person …

People used for recording of the data were wearing four tags (ankle left, ankle right, belt and chest). Each instance is a localization data for one of t…

classification, sequential, time-series, univariate

Life Sciences

Badges

Part of the problem in using an automated program to discover the unknown target function is to decide how to encode names such that the program can be us…

classification, text, univariate

Others

CNAE-9

This is a data set containing 1080 documents of free text business descriptions of Brazilian companies categorized into a subset of 9 categories cataloged…

classification, multivariate, text

Business

Indoor User Movement Predicti…

This dataset represents a real-life benchmark in the area of Ambient Assisted Living applications, as described in [1]. The binary classification task con…

classification, multivariate, sequential, time-series

Computer Science

Open University Learning Anal…

Open University Learning Analytics Dataset (OULAD) contains data about courses, students and their interactions with Virtual Learning Environment (VLE) fo…

classification, clustering, multivariate, regression, sequential, time-series

Computer Science

Machine Learning based ZZAlph…

1. Title of Database: Machine Learning based ZZAlpha Stock Recommendations 2. Sources: (a) Original owners of data: ZZAlpha Ltd., 4729 E. Sunrise #109,…

classification, sequential, time-series

Business

Data for Software Engineering…

The data can be used to try to predict student learning in SE teamwork based on observation of their team activity **** README FILE from the submitted d…

classification, sequential, time-series

Computer Science

UJI Pen Characters (Version 2)

We have created the UJIpenchars2 character database by collecting samples from 60 writers at two different sites in two phases: 1st phase, 11 writers, car…

classification, multivariate, sequential

Computer Science

Opinosis Opinion Review

This dataset contains sentences extracted from user reviews on a given topic. Example topics are performance of Toyota Camry and sound quality of ipod nan…

text

Computer Science

Bag of Words

For each text collection, D is the number of documents, W is the number of words in the vocabulary, and N is the total number of words in the collection (…

clustering, text

Others

YouTube Comedy Slam Preferenc…

YouTube Comedy Slam ([Web Link]) is a video discovery experiment running on YouTube's version of labs (called TestTube) for a few months in 2011 and 2012.…

classification, text

Computer Science

Legal Case Reports

This dataset contains Australian legal cases from the Federal Court of Australia (FCA). The cases were downloaded from AustLII ([Web Link]). We included a…

classification, text

Others

Miskolc IIS Hybrid IPS

The measurements were created to ease the development, comparison and evaluation of fingerprinting based hybrid indoor positioning methods. The measuremen…

causal-discovery, classification, clustering, text

Computer Science

Entree Chicago Recommendation…

This data records interactions with Entree Chicago restaurant recommendation system (originally [Web Link]) from September, 1996 to April, 1999. The data …

recommender, sequential, transactional

Others

Online Handwritten Assamese C…

A dataset of online handwritten assamese characters by collecting samples from 45 writers is created. Each writer contributed 52 basic characters, 10 nume…

classification, multivariate, sequential

Computer Science

SMS Spam Collection

This corpus has been collected from free or free for research sources at the Internet: -> A collection of 425 SMS spam messages was manually extracted fr…

classification, clustering, domain-theory, multivariate, text

Computer Science

UNIX User Data

Homepage

Description

Related Papers

Tags

Discussion

Related datasets