[python]赶集网二手房爬虫插件【可用任意扩展】

最近应一个老铁的要求，人家是搞房产的，所以就写了这个二手房的爬虫，因为初版，所以比较简单，有能力的老铁可用进行扩展。

import requests
import os
 
from bs4 import BeautifulSoup
 
 
 
class GanJi():
    """docstring for GanJi"""
 
    def __init__(self):
        super(GanJi, self).__init__()
 
    def get(self,url):
 
        user_agent = 'Mozilla/5.0 (Windows NT 6.3; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/52.0.2743.82 Safari/537.36'
        headers    = {'User-Agent':user_agent}
         
        webData    = requests.get(url + 'o1',headers=headers).text
        soup       = BeautifulSoup(webData,'lxml')
         
         
        sum        = soup.find('span',class_="num").text.replace("套","")
        ave        = int(sum) / 32
        forNum     = int(ave)
 
        if forNum < ave:
            forNum = forNum + 1
 
 
        for x in range(forNum):
            webData    = requests.get(url + 'o' + str(x + 1),headers=headers).text
            soup       = BeautifulSoup(webData,'lxml')
            find_list  = soup.find('div',class_="f-main-list").find_all('div',class_="f-list-item ershoufang-list")
 
            for dl in find_list:
                 
                print(dl.find('a',class_="js-title value title-font").text,end='|') # 名称
 
                # 中间 5 个信息
                tempDD = dl.find('dd',class_="dd-item size").find_all('span')
                for tempSpan in tempDD:
                    if not tempSpan.text == '' : 
                        print(tempSpan.text.replace("
", ""),end='|')
 
                 
                print(dl.find('span',class_="area").text.replace(" ","").replace("
",""),end='|') # 地址
                 
                print(dl.find('div',class_="price").text.replace(" ","").replace("
",""),end='|') # 价钱
                 
                print(dl.find('div',class_="time").text.replace(" ","").replace("
",""),end="|") # 平均
                 
                print("http://chaozhou.ganji.com" + dl['href'],end="|") # 地址
 
                print(str(x + 1))
 
if __name__ == '__main__':
    temp = GanJi()
    temp.get("http://chaozhou.ganji.com/fang5/xiangqiao/")

相关阅读:
MyBatis笔记----Mybatis3.4.2与spring4整合：增删查改
 MyBatis笔记----（2017年）最新的报错：Cannot find class [org.apache.commons.dbcp.BasicDataSource] for bean with name 'dataSource' defined in class path resource [com/ij34/mybatis/applicationContext.xml]; nested e
MyBatis笔记----报错：Error creating bean with name 'sqlSessionFactory' defined in class path resource [com/ij34/mybatis/applicationContext.xml]: Invocation of init method failed; nested exception is org.sp
MyBatis笔记----报错Exception in thread "main" org.apache.ibatis.binding.BindingException: Invalid bound statement (not found): com.ij34.model.UserMapper.selectUser
MyBatis笔记----报错：Exception in thread "main" org.apache.ibatis.binding.BindingException: Invalid bound statement (not found)解决方法
 MyBatis笔记----多表关联查询两种方式实现
 MyBatis笔记----MyBatis数据库表格数据修改更新的两种方法：XML与注解
 MyBatis笔记----MyBatis查询表全部的两种方法：XML与注解
 MyBatis笔记----MyBatis 入门经典的两个例子： XML 定义与注解定义
 springmvc复习笔记----文件上传multipartResolver
原文地址：https://www.cnblogs.com/68xi/p/9486957.html