BeautifulSoup 抓取网站url

  1 # -*- coding:utf-8 -*-
  2 import urlparse
  3 import urllib2
  4 from bs4 import BeautifulSoup
  5 
  6 url = "http://www.baidu.com"
  7 
  8 urls = [url] # stack of urls to scrape
  9 visited = [url] # historic record of urls
 10 
  1 # -*- coding:utf-8 -*-
  2 import urlparse
  3 import urllib2
  4 from bs4 import BeautifulSoup
  5 
  6 url = "http://www.baidu.com"
  7 
  8 urls = [url] # stack of urls to scrape
  9 visited = [url] # historic record of urls
 10 
 11 while len(urls) > 0:
 12     try:
 13         htmltext = urllib2.urlopen(urls[0]).read()
 14     except:
 15         print urls[0]
 16     soup = BeautifulSoup(htmltext,"html")
 17 
 18     urls.pop(0)
 19 
 20     for tag in soup.findAll("a", href=True):
 21         tag["href"] = urlparse.urljoin(url, tag["href"])
 22         if url in tag["href"] and tag["href"] not in visited:
 23             urls.append(tag["href"])
 24             visited.append(tag["href"])
 25 
 26     print len(urls)

相关阅读:
《软件安全系统设计最佳实践》课程大纲
《基于IPD流程的成本管理架构、方法和管理实践》培训大纲
《技术规划与路标开发实践》内训在芜湖天航集团成功举办!
年终总结，究竟该写什么？
Docker安装RabbitMQ
Ubuntu18安装docker
Ubuntu18.04安装MySQL
Windows常用CMD命令
ping不通域名的问题OR请求找不到主机-问题
JMeter 函数用法

原文地址：https://www.cnblogs.com/cuzz/p/BeautifulSoup.html