腳本之家服務(wù)器常用軟件

快捷導(dǎo)航

軟件下載

android MAC 驅(qū)動(dòng)下載字體下載 DLL

源碼下載

PHP ASP.NET ASP JSP

軟件編程

C# JAVA C 語言 Delphi Android

網(wǎng)絡(luò)編程

PHP ASP.NET ASP JavaScript

在線工具

CSS格式化 JS格式化 Html轉(zhuǎn)化為Js

數(shù)據(jù)庫

MYSQL MSSQL oracle DB2 MARIADB

CMS

PHPCMS DEDECMS 帝國CMS WordPress

常用工具

PHP開發(fā)工具 python Photoshop 必備軟件

Python爬蟲入門案例之回車桌面壁紙網(wǎng)美女圖片采集

更新時(shí)間：2021年10月15日 10:11:15 作者：松鼠愛吃餅干

讀萬卷書不如行萬里路，學(xué)的扎不扎實(shí)要通過實(shí)戰(zhàn)才能看出來，今天小編給大家?guī)硪粋€(gè)python爬蟲案例，采集回車桌面網(wǎng)站的美女圖片,大家可以在過程中查缺補(bǔ)漏，看看自己掌握程度怎么樣

知識(shí)點(diǎn)

requests
parsel
re
os

環(huán)境

python3.8
pycharm2021

目標(biāo)網(wǎng)址

https://mm.enterdesk.com/bizhi/63899-347866.html

【付費(fèi)VIP完整版】只要看了就能學(xué)會(huì)的教程，80集Python基礎(chǔ)入門視頻教學(xué)

點(diǎn)這里即可免費(fèi)在線觀看

注意: 在我們查看網(wǎng)頁源代碼的時(shí)候 (1. 控制臺(tái)為準(zhǔn) 2. 以右鍵查看網(wǎng)頁源代碼 3. 元素面板)

發(fā)送網(wǎng)絡(luò)請(qǐng)求
獲取網(wǎng)頁源代碼
提取想要的圖片鏈接 css樣式提取 xpath re正則表達(dá)式 bs4
替換所有的圖片鏈接換成大圖
保存圖片

爬蟲代碼

導(dǎo)入模塊

import requests     # 第三方庫 pip install requests
import parsel       # 第三方庫 pip install parsel
import os           # 新建文件夾

發(fā)送網(wǎng)絡(luò)請(qǐng)求

response = requests.get('https://mm.enterdesk.com/bizhi/64011-348522.html')

獲取網(wǎng)頁源代碼

data_html = response_1.text

提取每個(gè)相冊(cè)的詳情頁鏈接地址

selector_1 = parsel.Selector(data_html)
photo_url_list = selector_1.css('.egeli_pic_dl dd a::attr(href)').getall()
title_list = selector_1.css('.egeli_pic_dl dd a img::attr(title)').getall()
for photo_url, title in zip(photo_url_list, title_list):
    print(f'*****************正在爬取{title}*****************')
    response = requests.get(photo_url)
    # <Response [200]>: 請(qǐng)求成功的標(biāo)識(shí)
    selector = parsel.Selector(response.text)
    # 提取想要的圖片鏈接[第一個(gè)鏈接, 第二個(gè)鏈接,....]
    img_src_list = selector.css('.swiper-wrapper a img::attr(src)').getall()
    # 新建一個(gè)文件夾
    if not os.path.exists('img/' + title):
        os.mkdir('img/' + title)

替換所有的圖片鏈接換成大圖

for img_src in img_src_list:
    # 字符串的替換
    img_url = img_src.replace('_360_360', '_source')

保存圖片圖片名字

# 圖片 音頻 視頻 二進(jìn)制數(shù)據(jù)content
img_data = requests.get(img_url).content
# 圖片名稱 字符串分割
# 分割完之后 會(huì)給我們返回一個(gè)列表
img_title = img_url.split('/')[-1]
with open(f'img/{title}/{img_title}', mode='wb') as f:
    f.write(img_data)
print(img_title, '保存成功!!!')

翻頁

page_html = requests.get('https://mm.enterdesk.com/').text
counts = parsel.Selector(page_html).css('.wrap.no_a::attr(href)').get().split('/')[-1].split('.')[0]
for page in range(1, int(counts) + 1):
    print(f'------------------------------------正在爬取第{page}頁------------------------------------')
    發(fā)送網(wǎng)絡(luò)請(qǐng)求
    response_1 = requests.get(f'https://mm.enterdesk.com/{page}.html')