天天看点

python爬虫 requests.get()返回值与html网页不一致

写爬虫时,需要的html和用requests.get返回的html不一样导致后面用bs老出错

requests.get()获取不到正确的源代码HTML

# 1. 获取网页数据
url = 'https://movie.douban.com/top250'
headers = {
    'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.51 Safari/537.36'
}
response = requests.get(url, headers=headers)

# 2. 解析数据
soup = BeautifulSoup(response.text, 'lxml')
           

这个不行

# 指定要爬取的网站
    url = 'http://www.360doc.com/index.html?type=36&classid=19'
    soup = getsoup(url)
    print(soup)
    # 错了这么多,soup中竟没有
    imgList =soup.select('.c5_ul3>li') # 上下两标签内容 .class名>下级 标签
           

试了下下面的网址不行,更换headers一样不同:

#获取网页数据
url = 'https://movie.douban.com/tv'
headers = {
    'user-agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/99.0.4844.51 Safari/537.36'
}
response = requests.get(url, headers=headers)

# 2. 解析数据
soup = BeautifulSoup(response.text, 'lxml')
           

这个库,没看出来为什么,有的网页可以,有的却是错的