Node.js 爬蟲初探

本文轉載自查看原文 2015-12-06 22:11 1964 Node.js

前言

在學習慕課網視頻和Cnode新手入門接觸到爬蟲，說是爬蟲初探，其實並沒有用到爬蟲相關第三方類庫，主要用了node.js基礎模塊http、網頁分析工具cherrio。使用http直接獲取url路徑對應網頁資源，然后使用cherrio分析。這里我主要是把慕課網教學視頻提供的案例自己敲了一邊，加深理解。在coding的過程中，我第一次把jq獲取后的對象直接用forEach遍歷，直接報錯，是因為jq沒有對應的這個方法，只有js數組可以調用。

知識點

①：superagent抓去網頁工具。我暫時未用到。

②：cherrio 網頁分析工具，你可以理解其為服務端的jQuery，因為語法都一樣。

效果圖

1、抓取整個網頁

2、分析后的數據，我這里是以慕課網提供的示例為案例實現的例子。

爬蟲初探源碼分析

var http=require('http');
var cheerio=require('cheerio');

var url='http://www.imooc.com/learn/348';

/****************************
打印得到的數據結構
[{
	chapterTitle:'',
	videos:[{
		title:'',
		id:''
	}]
}]
********************************/
function printCourseInfo(courseData){
	courseData.forEach(function(item){
		var chapterTitle=item.chapterTitle;
		console.log(chapterTitle+'\n');
		item.videos.forEach(function(video){
			console.log(' 【'+video.id+'】'+video.title+'\n');
		})
	});
}


/*************
分析從網頁里抓取到的數據
**************/
function filterChapter(html){
	var courseData=[];

	var $=cheerio.load(html);
	var chapters=$('.chapter');
	chapters.each(function(item){
		var chapter=$(this);
		var chapterTitle=chapter.find('strong').text(); //找到章節標題
		var videos=chapter.find('.video').children('li');

		var chapterData={
			chapterTitle:chapterTitle,
			videos:[]
		};

		videos.each(function(item){
			var video=$(this).find('.studyvideo'); 
			var title=video.text();
			var id=video.attr('href').split('/video')[1];

			chapterData.videos.push({
				title:title,
				id:id
			})
		})

		courseData.push(chapterData);
	});

    return courseData;
}

http.get(url,function(res){
	var html='';

	res.on('data',function(data){
		html+=data;
	})

	res.on('end',function(){
		var courseData=filterChapter(html);
		printCourseInfo(courseData);
	})
}).on('error',function(){
	console.log('獲取課程數據出錯');
})

參考資料

https://github.com/alsotang/node-lessons/tree/master/lesson3

http://www.imooc.com/video/7965

免責聲明！

本站轉載的文章為個人學習借鑒使用，本站對版權不負任何法律責任。如果侵犯了您的隱私權益，請聯系本站郵箱yoyou2525@163.com刪除。

猜您在找 Node.js源碼初探~我很好奇 node.js之cluster集群初探基於Node.js的爬蟲工具 – Node Crawler node.js 爬蟲動態代理ip Node.js大眾點評爬蟲 Node.js 實現簡單小說爬蟲 Node.js爬蟲--網頁請求模塊 [轉] Node.js 服務端實踐之 GraphQL 初探初探 Node.js 框架：eggjs （環境搭配篇）基於node.js的爬蟲框架 node-crawler簡單嘗試