# geccoDemo **Repository Path**: lh880/geccoDemo ## Basic Information - **Project Name**: geccoDemo - **Description**: geccoDemo java 爬虫 - **Primary Language**: Java - **License**: Not specified - **Default Branch**: master - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 6 - **Created**: 2019-12-04 - **Last Updated**: 2020-12-19 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README ##JAVA 轻量级爬虫框架 Gecco 使用Demo 主要使用Gecco抓取新闻网站所有新闻 ##代码启动 ``` public static void main(String[] rags) { List urls=new ArrayList(); urls.add( new HttpGetRequest("http://zj.zjol.com.cn/home.html?pageIndex=1&pageSize=100")); urls.add( new HttpGetRequest("http://zj.zjol.com.cn/politics.html?pageIndex=1&pageSize=100")); urls.add( new HttpGetRequest("http://zj.zjol.com.cn/original.html?pageIndex=1&pageSize=100")); urls.add( new HttpGetRequest("http://zj.zjol.com.cn/life.html?pageIndex=1&pageSize=100")); urls.add( new HttpGetRequest("http://zj.zjol.com.cn/vision.html?pageIndex=1&pageSize=100")); urls.add( new HttpGetRequest("http://zj.zjol.com.cn/huatuxia.html?pageIndex=1&pageSize=100")); GeccoEngine.create() //工程的包路径 .classpath("com.zhaochao.gecco.zj") //开始抓取的页面地址 .start(urls) //开启几个爬虫线程 .thread(10) //单个爬虫每次抓取完一个请求后的间隔时间 .interval(10) //使用pc端userAgent .mobile(false) //开始运行 .run(); } ``` ##抓取分页 ``` if(page<=maxPage){ System.out.println("type=" + type+ " now page = "+page); String nextUrl = "http://zj.zjol.com.cn/home.html?pageIndex="+page+"&pageSize=100"; //抓取下一页 SchedulerContext.into(request.subRequest(nextUrl)); } ``` ##抓取祥情页 ``` for (HrefBean bean:zjNewsGeccoList.getNewList()){ //进入祥情页面抓取 SchedulerContext.into(request.subRequest("http://zj.zjol.com.cn"+bean.getUrl())); } ``` ##通过注解及CSS选择器抓取html节点 ``` @Text @HtmlField(cssPath = "#headline") private String title ; @Text @HtmlField(cssPath = "#content > div > div.news_con > div.news-content > div:nth-child(1) > div > p.go-left.post-time.c-gray") private String createTime; ```