在用R語言做數據分析的時候,我們經常需要一些數據做實驗,尤其當學會一個新的方法或者算法的時候,就想趕快編程實現一下。那么,問題來了,用什么數據好呢,什么樣的數據適合做這種實驗呢?
好在R語言提供了很多的基本數據集,這些基本數據集是可以直接加載、可以完成幾乎所有的數據分析任務模擬數據的。這都是些哪些種類的數據呢?
之前我們安裝了 ggplot2 會使用其默認數據集,其實R語言自帶也是有數據集的。
data(package = 'ggplot2')
R自帶基本數據集都在R包datasets里面,可以執行命令:
data(package = 'datasets')
就會datasets包帶的出現所有基本基本數據集:
Data sets in package ‘datasets’: AirPassengers Monthly Airline Passenger Numbers 1949-1960 BJsales Sales Data with Leading Indicator BJsales.lead (BJsales) Sales Data with Leading Indicator BOD Biochemical Oxygen Demand CO2 Carbon Dioxide Uptake in Grass Plants ChickWeight Weight versus age of chicks on different diets DNase Elisa assay of DNase EuStockMarkets Daily Closing Prices of Major European Stock Indices, 1991-1998 Formaldehyde Determination of Formaldehyde HairEyeColor Hair and Eye Color of Statistics Students Harman23.cor Harman Example 2.3 Harman74.cor Harman Example 7.4 Indometh Pharmacokinetics of Indomethacin InsectSprays Effectiveness of Insect Sprays JohnsonJohnson Quarterly Earnings per Johnson & Johnson Share LakeHuron Level of Lake Huron 1875-1972 LifeCycleSavings Intercountry Life-Cycle Savings Data Loblolly Growth of Loblolly pine trees Nile Flow of the River Nile Orange Growth of Orange Trees OrchardSprays Potency of Orchard Sprays PlantGrowth Results from an Experiment on Plant Growth Puromycin Reaction Velocity of an Enzymatic Reaction Seatbelts Road Casualties in Great Britain 1969-84 Theoph Pharmacokinetics of Theophylline Titanic Survival of passengers on the Titanic ToothGrowth The Effect of Vitamin C on Tooth Growth in Guinea Pigs UCBAdmissions Student Admissions at UC Berkeley UKDriverDeaths Road Casualties in Great Britain 1969-84 UKgas UK Quarterly Gas Consumption USAccDeaths Accidental Deaths in the US 1973-1978 USArrests Violent Crime Rates by US State USJudgeRatings Lawyers' Ratings of State Judges in the US Superior Court USPersonalExpenditure Personal Expenditure Data UScitiesD Distances Between European Cities and Between US Cities VADeaths Death Rates in Virginia (1940) WWWusage Internet Usage per Minute WorldPhones The World's Telephones ability.cov Ability and Intelligence Tests airmiles Passenger Miles on Commercial US Airlines, 1937-1960 airquality New York Air Quality Measurements anscombe Anscombe's Quartet of 'Identical' Simple Linear Regressions attenu The Joyner-Boore Attenuation Data attitude The Chatterjee-Price Attitude Data austres Quarterly Time Series of the Number of Australian Residents beaver1 (beavers) Body Temperature Series of Two Beavers beaver2 (beavers) Body Temperature Series of Two Beavers cars Speed and Stopping Distances of Cars chickwts Chicken Weights by Feed Type co2 Mauna Loa Atmospheric CO2 Concentration crimtab Student's 3000 Criminals Data discoveries Yearly Numbers of Important Discoveries esoph Smoking, Alcohol and (O)esophageal Cancer euro Conversion Rates of Euro Currencies euro.cross (euro) Conversion Rates of Euro Currencies eurodist Distances Between European Cities and Between US Cities faithful Old Faithful Geyser Data fdeaths (UKLungDeaths) Monthly Deaths from Lung Diseases in the UK freeny Freeny's Revenue Data freeny.x (freeny) Freeny's Revenue Data freeny.y (freeny) Freeny's Revenue Data infert Infertility after Spontaneous and Induced Abortion iris Edgar Anderson's Iris Data iris3 Edgar Anderson's Iris Data islands Areas of the World's Major Landmasses ldeaths (UKLungDeaths) Monthly Deaths from Lung Diseases in the UK lh Luteinizing Hormone in Blood Samples longley Longley's Economic Regression Data lynx Annual Canadian Lynx trappings 1821-1934 mdeaths (UKLungDeaths) Monthly Deaths from Lung Diseases in the UK morley Michelson Speed of Light Data mtcars Motor Trend Car Road Tests nhtemp Average Yearly Temperatures in New Haven nottem Average Monthly Temperatures at Nottingham, 1920-1939 npk Classical N, P, K Factorial Experiment occupationalStatus Occupational Status of Fathers and their Sons precip Annual Precipitation in US Cities presidents Quarterly Approval Ratings of US Presidents pressure Vapor Pressure of Mercury as a Function of Temperature quakes Locations of Earthquakes off Fiji randu Random Numbers from Congruential Generator RANDU rivers Lengths of Major North American Rivers rock Measurements on Petroleum Rock Samples sleep Student's Sleep Data stack.loss (stackloss) Brownlee's Stack Loss Plant Data stack.x (stackloss) Brownlee's Stack Loss Plant Data stackloss Brownlee's Stack Loss Plant Data state.abb (state) US State Facts and Figures state.area (state) US State Facts and Figures state.center (state) US State Facts and Figures state.division (state) US State Facts and Figures state.name (state) US State Facts and Figures state.region (state) US State Facts and Figures state.x77 (state) US State Facts and Figures sunspot.month Monthly Sunspot Data, from 1749 to "Present" sunspot.year Yearly Sunspot Data, 1700-1988 sunspots Monthly Sunspot Numbers, 1749-1983 swiss Swiss Fertility and Socioeconomic Indicators (1888) Data treering Yearly Treering Data, -6000-1979 trees Diameter, Height and Volume for Black Cherry Trees uspop Populations Recorded by the US Census volcano Topographic Information on Auckland's Maunga Whau Volcano warpbreaks The Number of Breaks in Yarn during Weaving women Average Heights and Weights for American Women
具體翻譯一下各個數據集基本情況
向量
euro #歐元匯率,長度為11,每個元素都有命名
landmasses #48個陸地的面積,每個都有命名
precip #長度為70的命名向量
rivers #北美141條河流長度
state.abb #美國50個州的雙字母縮寫
state.area #美國50個州的面積
state.name #美國50個州的全稱
因子
state.division #美國50個州的分類,9個類別
state.region #美國50個州的地理分類
矩陣、數組
euro.cross #11種貨幣的匯率矩陣
freeny.x #每個季度影響收入四個因素的記錄
state.x77 #美國50個州的八個指標
USPersonalExpenditure #5個年份在5個消費方向的數據
VADeaths #1940年弗吉尼亞州死亡率(每千人)
volcano #某火山區的地理信息(10米×10米的網格)
WorldPhones #8個區域在7個年份的電話總數
iris3 #3種鳶尾花形態數據
Titanic #泰坦尼克乘員統計
UCBAdmissions #伯克利分校1973年院系、錄取和性別的頻數
crimtab #3000個男性罪犯左手中指長度和身高關系
HairEyeColor #592人頭發顏色、眼睛顏色和性別的頻數
occupationalStatus #英國男性父子職業聯系
類矩陣
eurodist #歐洲12個城市的距離矩陣,只有下三角部分
Harman23.cor #305個女孩八個形態指標的相關系數矩陣
Harman74.cor #145個兒童24個心理指標的相關系數矩陣
數據框
airquality #紐約1973年5-9月每日空氣質量
anscombe #四組x-y數據,雖有相似的統計量,但實際數據差別較大
attenu #多個觀測站對加利福尼亞23次地震的觀測數據
attitude #30個部門在七個方面的調查結果,調查結果是同一部門35個職員贊成的百分比
beaver1 #一只海狸每10分鍾的體溫數據,共114條數據
beaver2 #另一只海狸每10分鍾的體溫數據,共100條數據
BOD #隨水質的提高,生化反應對氧的需求(mg/l)隨時間(天)的變化
cars #1920年代汽車速度對剎車距離的影響
chickwts #不同飲食種類對小雞生長速度的影響
esoph #法國的一個食管癌病例對照研究
faithful #一個間歇泉的爆發時間和持續時間
Formaldehyde #兩種方法測定甲醛濃度時分光光度計的讀數
Freeny #每季度收入和其他四因素的記錄
dating from #配對的病例對照數據,用於條件logistic回歸
InsectSprays #使用不同殺蟲劑時昆蟲數目
iris #3種鳶尾花形態數據
LifeCycleSavings #50個國家的存款率
longley #強共線性的宏觀經濟數據
morley #光速測量試驗數據
mtcars #32輛汽車在11個指標上的數據
OrchardSprays #使用拉丁方設計研究不同噴霧劑對蜜蜂的影響
PlantGrowth #三種處理方式對植物產量的影響
pressure #溫度和氣壓
Puromycin #兩種細胞中輔因子濃度對酶促反應的影響
quakes #1000次地震觀測數據(震級>4)
randu #在VMS1.5中使用FORTRAN中的RANDU三個一組生成隨機數字,共400組。
rock #48塊石頭的形態數據
sleep #兩葯物的催眠效果
stackloss #化工廠將氨轉為硝酸的數據
swiss #瑞士生育率和社會經濟指標
ToothGrowth #VC劑量和攝入方式對豚鼠牙齒的影響
trees #樹木形態指標
USArrests #美國50個州的四個犯罪率指標
USJudgeRatings #43名律師的12個評價指標
warpbreaks #織布機異常數據
women #15名女性的身高和體重
列表
state.center #美國50個州中心的經度和緯度
類數據框
ChickWeight #飲食對雞生長的影響
CO2 #耐寒植物CO2攝取的差異
DNase #若干次試驗中,DNase濃度和光密度的關系
Indometh #某葯物的葯物動力學數據
Loblolly #火炬松的高度、年齡和種源
Orange #桔子樹生長數據
Theoph #茶鹼葯動學數據
時間序列數據
airmiles #美國1937-1960年客運里程營收(實際售出機位乘以飛行哩數)
AirPassengers #Box & Jenkins航空公司1949-1960年每月國際航線乘客數
austres #澳大利亞1971-1994每季度人口數(以千為單位)
BJsales #有關銷售的一個時間序列
BJsales.lead #前一指標的先行指標(leading indicator)
co2 #1959-1997年每月大氣co2濃度(ppm)
discoveries #1860-1959年每年巨大發現或發明的個數
ldeaths #1974-1979年英國每月支氣管炎、肺氣腫和哮喘的死亡率
fdeaths #前述死亡率的女性部分
mdeaths #前述死亡率的男性部分
freeny.y #每季度收入
JohnsonJohnson #1960-1980年每季度Johnson & Johnson股票的紅利
LakeHuron #1875-1972年某一湖泊水位的記錄
lh #黃體生成素水平,10分鍾測量一次
lynx #1821-1934年加拿大猞猁數據
nhtemp #1912-1971年每年平均溫度
Nile #1871-1970尼羅河流量
nottem #1920-1939每月大氣溫度
presidents #1945-1974年每季度美國總統支持率
UKDriverDeaths #1969-1984年每月英國司機死亡或嚴重傷害的數目
sunspot.month #1749-1997每月太陽黑子數
sunspot.year #1700-1988每年太陽黑子數
sunspots #1749-1983每月太陽黑子數
treering #歸一化的樹木年輪數據
UKgas #1960-1986每月英國天然氣消耗
USAccDeaths #1973-1978美國每月意外死亡人數
uspop #1790–1970美國每十年一次的人口總數(百萬為單位)
WWWusage #每分鍾網絡連接數
Seatbelts #多變量時間序列。和UKDriverDeaths時間段相同,反映更多因素。
EuStockMarkets #多變量時間序列。歐洲股市四個主要指標的每個工作日記錄,共1860條記錄。
當然了,如果我們想知道任意一個包自帶有哪些數據集,可以有命令:> data(package = '包名')
data(package = 'ggplot2')
Data sets in package ‘ggplot2’: diamonds Prices of over 50,000 round cut diamonds economics US economic time series economics_long US economic time series faithfuld 2d density estimate of Old Faithful data luv_colours 'colors()' in Luv space midwest Midwest demographics mpg Fuel economy data from 1999 to 2008 for 38 popular models of cars msleep An updated and expanded version of the mammals sleep dataset presidential Terms of 11 presidents from Eisenhower to Obama seals Vector field of seal movements txhousing Housing sales in TX