EADST

Extract Webpage Information with Python

Here is the python program to extract webpage information with BeautifulSoup and save the data in a CSV file.

from bs4 import BeautifulSoup
import urllib.request
import pandas as pd

url = 'file:///Users/xd/Desktop/ieee/Region_5_Student_Branch_Counselors_and_Chairs.htm'
save_file = 'ieee_info_1'
html = urllib.request.urlopen(url).read()

soup = BeautifulSoup(html, "html.parser")

universities = soup.find_all('div', class_='spoName bullet pad-t15')
people = soup.find_all('div', class_='roster-results')

for u, p in zip(universities, people):
    info = p.find_all('p')
    university = u.get_text()
    name = info[0].get_text()
    if name == 'Position Vacant':
        continue
    title = info[2].get_text()
    address = info[3].get_text() + ', ' + info[4].get_text()
    email = info[-1].get_text()[7:]

    content = [[university, name, title, address, email]]
    list_name = ['university', 'name', 'title', 'address', 'email']
    data = pd.DataFrame(columns=list_name, data=content)
    data.to_csv("{}.csv".format(save_file), mode='a', index=False, header=False, encoding='utf-8')
About Me
XD
Goals determine what you are going to be.
Category
标签云
Django VGG-16 Pandas Pytorch 搞笑 云服务器 hf 公式 Miniforge News Distillation Video Firewall Safetensors GPT4 论文速读 ResNet-50 git-lfs 论文 Plotly Base64 logger Jetson FP8 SQLite Permission Magnet 腾讯云 API网关 DeepStream Quantization Web PIP Website SPIE v0.dev TensorFlow HuggingFace Zip COCO NLTK YOLO Food ms-swift LLM XML InvalidArgumentError Linux UNIX CV Card Qwen2 VSCode tqdm Data AI Sklearn 音频 阿里云 顶会 Mixtral Claude Markdown RL CC TTS LLAMA Github CEIR C++ SQL UI Conda MD5 Hilton 证件照 Windows Ptyhon Breakpoint CUDA Plate 继承 EXCEL PDB IndexTTS2 RAR Agent Proxy Paddle 关于博主 Math Bitcoin WAN Bert API Heatmap NameSilo Interview Template Land SVR PyCharm Freesound Qwen Augmentation Attention OpenCV 净利润 签证 NLP scipy Tiktoken printf Google QWEN Crawler FP32 FP16 Python Color 飞书 Bipartite TSV Michelin Qwen2.5 Search VPN PyTorch Input GIT Algorithm transformers ChatGPT Llama Clash icon 图形思考法 Shortcut BeautifulSoup Ubuntu TensorRT LoRA WebCrawler Jupyter BF16 第一性原理 OpenAI Image2Text Hungarian CAM 递归学习法 Vim Statistics FlashAttention LeetCode Anaconda Paper Rebuttal 多线程 mmap Cloudreve Nginx uwsgi Harness Review Translation SAM Disk 图标 Streamlit Tensor Pickle Hotel DeepSeek Docker Transformers Gemma Quantize tar v2ray Diagram Knowledge Password Domain GPTQ GoogLeNet Jev Use FastAPI CSV Datetime 版权 Logo llama.cpp BTC Vmess Bin Animate RGB FP64 Numpy Random Git LaTeX Tracking torchinfo XGBoost Pillow ModelScope ONNX HaggingFace Baidu GGML Dataset 强化学习 CTC 财报 uWSGI 算法题 CLAP 域名 JSON 多进程 OCR PDF diffusers 报税 Excel git
站点统计

本站现有博文337篇,共被浏览957675次

本站已经建立2669天!

热门文章
文章归档
回到顶部