python爬虫——使用selenium+chrome options爬取站长素材页面源码

2023-08-05 23:25:12

一.站长素材

1.需要爬取的内容

2.代码

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
import time
# webdriver 路径
path = r'E:\chromedriver_win32\chromedriver.exe'
# 创建无界面浏览器
chrome_options = Options()
chrome_options.add_argument("--headless")
browser = webdriver.Chrome(executable_path=path, options=chrome_options)

# 站长素材高清图片-科技图片url
url = 'http://sc.chinaz.com/tupian/kejitupian.html'
browser.get(url)
time.sleep(3)
# 第一次保存html代码
with open('kejitupian1.html', 'w', encoding='utf8') as fp:
    fp.write(browser.page_source)
# 滚动，执行js
js = 'window.scrollTo(0,document.body.scrollHeight)'
browser.execute_script(js)
time.sleep(3)
# 第二次保存html代码
with open('kejitupian2.html', 'w', encoding='utf8') as fp:
    fp.write(browser.page_source)

# 关闭浏览器
browser.quit()

3.结果对比

第一次抓取：

python爬虫——使用selenium+chrome options爬取站长素材页面源码

第二次抓取：

python爬虫——使用selenium+chrome options爬取站长素材页面源码

python爬虫——使用selenium+chrome options爬取站长素材页面源码

一.站长素材

1.需要爬取的内容

2.代码

3.结果对比

继续阅读

Python爬虫之网站超清图片爬取(2021.3.29)

Python入门级爬取百度百科词条

16Python爬虫---Scrapy常用命令

Python爬虫基本库的使用第二章基本库的使用

Python爬虫（四）lxml、xpath安装模块导入查找节点属性查找 @ 符号使用谓语选取未知节点获取文本和属性

爬虫学习之04-request模块获取糗事百科一张热图

python3下用selenium库和chrome的headless模式实现网页抓取（注释中有用phantomJS的小段代码）

【Python爬虫案例学习19】多进程爬取某图片网站

python爬虫实战之爬取成语大全

【爬取百度首页】-将整个html源码保存-headers使用一、网页分析二、代码实现与步骤三、结果分析

爬取百度贴吧

爬取猫眼电影--静态网页反爬与多线程/多进程爬取网页解析爬取代码多线程与多进程

requests模块进行人人网模拟登陆

2023爬虫学习笔记 -- 多线程操作

Python爬虫学习（1）

Boss直聘Python爬虫实战