python技巧 - 筛选 Dataframe 数据的 10 种方法
liuian 2024-12-20 17:18 37 浏览
python工具包pandas提供了一种存储和处理的数据类型 dataframe,和数据库中的数据表很相似,这种数据类型也提供了query()查询数据的方法。从dataframe中筛选数据是经常用的数据操作方法,还介绍了 iloc 和 loc 的区别。本文把这些常用的方法汇集在一起,供学习者参考。
数据准备
import pandas as pd
df = pd.read_csv("flights_short.csv", usecols=range(1,17))
# 共111条数据
某国航空公司的飞行数据,格式和数据见本文末尾。
(01) 使用列值筛选数据
查找 飞机始发机场为“JFK” 且航空公司代号为“B6”的飞行记录,
newdf = df[(df.origin == "JFK") & (df.carrier == "B6")]
print(len(newdf))
print(newdf[0:10])
输出结果如下:
(02) 使用Query()函数
可以使用dataframe提供的函数query()查找,
newdf = df.query('origin == "JFK" & carrier == "B6"')
(03) 使用 loc 函数
newdf = df.loc[(df.origin == "JFK") & (df.carrier == "B6")]
(方法02和03的结果和方法01的结果一样)
(04) 使用行位置和列位置筛选
df.iloc[:5,] #找出前5行
df.iloc[1:5,] #找出第2行到第5行
df.iloc[5,0] #找出第6行第1列的值
df.iloc[1:5,0] #找出第2行到第5行的第1列
df.iloc[1:5,:5] #找出第2行到第5行的第5列
df.iloc[2:7,1:3] #找出第3到7行, 第2列到第3列
(05) 使用行位置和列名称筛选数据
df.loc[df.index[0:5],["origin","dest"]]
(06) 查找某一列的多个值
newdf = df.loc[(df.origin == "JFK") | (df.origin == "LGA")]
# 或者
newdf = df[df.origin.isin(["JFK", "LGA"])]
(07) 按指定条件搜索行数据
newdf = df.loc[(df.origin != "JFK") & (df.carrier == "B6")]
(08) 找出某一列中不重复的值
newpd.unique(newdf.origin)
结果为:['LGA', 'EWR']
(09) 找出不满足某个条件的值
newdf = df[df.origin.notnull()]
(10) 查找dataframe中的字符串
import pandas as pd
df = pd.DataFrame({"var1": ["AA_2", "B_1", "C_2", "A_2"]})
df
运行结果如下:
var1
0 AA_2
1 B_1
2 C_2
3 A_2
查找以A开头的字符串,
df[df['var1'].str[0] == 'A']
查找长度大于3的字符串,
df[df['var1'].str.len()>3]
查找包括字母A或B的字符串,
df[df['var1'].str.contains('A|B')]
筛选数据时,如何处理列名称中的空格
df.rename(columns={'var1':'var 1'}, inplace = True)
df
运行结果如下图所示:
loc 和 iloc 之间的区别
import numpy as np
x = pd.DataFrame({"col1" : np.arange(1,20,2)}, index=[9,8,7,6,0, 1, 2, 3, 4, 5])
结果如下图:
使用 iloc[0:5] 的结果:
使用 loc[0:5] 的结果:
其中,iloc 的结果中使用指定的索引标识,而loc则是默认的序列值。
附:本文实例中用到的数据及其格式:
"","year","month","day","dep_time","dep_delay","arr_time","arr_delay","carrier","tailnum","flight","origin","dest","air_time","distance","hour","minute"
"1",2013,1,1,517,2,830,11,"UA","N14228",1545,"EWR","IAH",227,1400,5,17
"2",2013,1,1,533,4,850,20,"UA","N24211",1714,"LGA","IAH",227,1416,5,33
"3",2013,1,1,542,2,923,33,"AA","N619AA",1141,"JFK","MIA",160,1089,5,42
"4",2013,1,1,544,-1,1004,-18,"B6","N804JB",725,"JFK","BQN",183,1576,5,44
"5",2013,1,1,554,-6,812,-25,"DL","N668DN",461,"LGA","ATL",116,762,5,54
"6",2013,1,1,554,-4,740,12,"UA","N39463",1696,"EWR","ORD",150,719,5,54
"7",2013,1,1,555,-5,913,19,"B6","N516JB",507,"EWR","FLL",158,1065,5,55
"8",2013,1,1,557,-3,709,-14,"EV","N829AS",5708,"LGA","IAD",53,229,5,57
"9",2013,1,1,557,-3,838,-8,"B6","N593JB",79,"JFK","MCO",140,944,5,57
"10",2013,1,1,558,-2,753,8,"AA","N3ALAA",301,"LGA","ORD",138,733,5,58
"11",2013,1,1,558,-2,849,-2,"B6","N793JB",49,"JFK","PBI",149,1028,5,58
"12",2013,1,1,558,-2,853,-3,"B6","N657JB",71,"JFK","TPA",158,1005,5,58
"13",2013,1,1,558,-2,924,7,"UA","N29129",194,"JFK","LAX",345,2475,5,58
"14",2013,1,1,558,-2,923,-14,"UA","N53441",1124,"EWR","SFO",361,2565,5,58
"15",2013,1,1,559,-1,941,31,"AA","N3DUAA",707,"LGA","DFW",257,1389,5,59
"16",2013,1,1,559,0,702,-4,"B6","N708JB",1806,"JFK","BOS",44,187,5,59
"17",2013,1,1,559,-1,854,-8,"UA","N76515",1187,"EWR","LAS",337,2227,5,59
"18",2013,1,1,600,0,851,-7,"B6","N595JB",371,"LGA","FLL",152,1076,6,0
"19",2013,1,1,600,0,837,12,"MQ","N542MQ",4650,"LGA","ATL",134,762,6,0
"20",2013,1,1,601,1,844,-6,"B6","N644JB",343,"EWR","PBI",147,1023,6,1
"21",2013,1,1,602,-8,812,-8,"DL","N971DL",1919,"LGA","MSP",170,1020,6,2
"22",2013,1,1,602,-3,821,16,"MQ","N730MQ",4401,"LGA","DTW",105,502,6,2
"23",2013,1,1,606,-4,858,-12,"AA","N633AA",1895,"EWR","MIA",152,1085,6,6
"24",2013,1,1,606,-4,837,-8,"DL","N3739P",1743,"JFK","ATL",128,760,6,6
"25",2013,1,1,607,0,858,-17,"UA","N53442",1077,"EWR","MIA",157,1085,6,7
"26",2013,1,1,608,8,807,32,"MQ","N9EAMQ",3768,"EWR","ORD",139,719,6,8
"27",2013,1,1,611,11,945,14,"UA","N532UA",303,"JFK","SFO",366,2586,6,11
"28",2013,1,1,613,3,925,4,"B6","N635JB",135,"JFK","RSW",175,1074,6,13
"29",2013,1,1,615,0,1039,-21,"B6","N794JB",709,"JFK","SJU",182,1598,6,15
"30",2013,1,1,615,0,833,-9,"DL","N326NB",575,"EWR","ATL",120,746,6,15
"31",2013,1,1,622,-8,1017,3,"US","N807AW",245,"EWR","PHX",342,2133,6,22
"32",2013,1,1,623,13,920,5,"AA","N3EMAA",1837,"LGA","MIA",153,1096,6,23
"33",2013,1,1,623,-4,933,1,"UA","N459UA",496,"LGA","IAH",229,1416,6,23
"34",2013,1,1,624,-6,909,29,"EV","N11107",4626,"EWR","MSP",190,1008,6,24
"35",2013,1,1,624,-6,840,10,"MQ","N518MQ",4599,"LGA","MSP",166,1020,6,24
"36",2013,1,1,627,-3,1018,0,"US","N535UW",27,"JFK","PHX",330,2153,6,27
"37",2013,1,1,628,-2,1137,-3,"AA","N3BAAA",413,"JFK","SJU",192,1598,6,28
"38",2013,1,1,628,-2,1016,29,"UA","N33289",1665,"EWR","LAX",366,2454,6,28
"39",2013,1,1,629,-1,824,14,"AA","N3CYAA",303,"LGA","ORD",140,733,6,29
"40",2013,1,1,629,-1,721,-19,"WN","N273WN",4646,"LGA","BWI",40,185,6,29
"41",2013,1,1,629,-1,824,-9,"US","N426US",1019,"EWR","CLT",91,529,6,29
"42",2013,1,1,632,24,740,12,"EV","N13553",4144,"EWR","IAD",52,212,6,32
"43",2013,1,1,635,0,1028,48,"AA","N3GKAA",711,"LGA","DFW",248,1389,6,35
"44",2013,1,1,637,-8,930,-5,"B6","N709JB",389,"LGA","MCO",144,950,6,37
"45",2013,1,1,639,-1,739,-10,"B6","N805JB",1002,"JFK","BOS",41,187,6,39
"46",2013,1,1,643,-3,922,-18,"UA","N497UA",556,"EWR","PBI",146,1023,6,43
"47",2013,1,1,643,-2,837,-11,"US","N178US",926,"EWR","CLT",91,529,6,43
"48",2013,1,1,644,8,931,-9,"UA","N75435",1701,"EWR","FLL",151,1065,6,44
"49",2013,1,1,645,-2,815,5,"B6","N796JB",102,"JFK","BUF",63,301,6,45
"50",2013,1,1,646,1,910,-6,"UA","N569UA",883,"LGA","DEN",243,1620,6,46
"51",2013,1,1,646,1,1023,-7,"UA","N38727",1496,"EWR","SNA",380,2434,6,46
"52",2013,1,1,651,-4,936,-6,"B6","N558JB",203,"JFK","LAS",323,2248,6,51
"53",2013,1,1,652,-3,932,11,"B6","N178JB",117,"JFK","MSY",191,1182,6,52
"54",2013,1,1,653,-7,936,-33,"DL","N327NW",1383,"LGA","PBI",149,1035,6,53
"55",2013,1,1,655,0,1021,-9,"DL","N3763D",1415,"JFK","SLC",294,1990,6,55
"56",2013,1,1,655,-5,1037,-8,"DL","N705TW",1865,"JFK","SFO",362,2586,6,55
"57",2013,1,1,655,-5,1002,-18,"DL","N997DL",2003,"LGA","MIA",161,1096,6,55
"58",2013,1,1,656,-4,854,4,"AA","N4WNAA",305,"LGA","ORD",143,733,6,56
"59",2013,1,1,656,-3,949,-10,"AA","N5FMAA",1815,"JFK","MCO",142,944,6,56
"60",2013,1,1,656,-9,1007,27,"MQ","N722MQ",4534,"LGA","XNA",233,1147,6,56
"61",2013,1,1,656,-4,948,-23,"UA","N24212",1115,"EWR","TPA",156,997,6,56
"62",2013,1,1,657,-3,959,-14,"DL","N318NB",1879,"LGA","FLL",164,1076,6,57
"63",2013,1,1,658,-2,944,5,"DL","N6703D",1547,"LGA","ATL",126,762,6,58
"64",2013,1,1,658,-2,1027,2,"VX","N627VA",399,"JFK","LAX",361,2475,6,58
"65",2013,1,1,659,-1,1008,-7,"AA","N3EKAA",2279,"LGA","MIA",159,1096,6,59
"66",2013,1,1,659,-1,1008,1,"B6","N646JB",981,"JFK","FLL",156,1069,6,59
"67",2013,1,1,659,-6,907,-6,"DL","N998DL",831,"LGA","DTW",105,502,6,59
"68",2013,1,1,659,-1,959,-9,"UA","N838UA",960,"EWR","RSW",164,1068,6,59
"69",2013,1,1,701,1,1123,-31,"UA","N77296",1203,"EWR","SJU",188,1608,7,1
"70",2013,1,1,702,2,1058,44,"B6","N779JB",671,"JFK","LAX",381,2475,7,2
"71",2013,1,1,709,9,852,20,"UA","N26226",1092,"LGA","ORD",135,733,7,9
"72",2013,1,1,711,-4,1151,-15,"B6","N651JB",715,"JFK","SJU",190,1598,7,11
"73",2013,1,1,712,-3,1023,-12,"AA","N3ETAA",825,"JFK","FLL",159,1069,7,12
"74",2013,1,1,715,2,911,21,"UA","N841UA",544,"EWR","ORD",156,719,7,15
"75",2013,1,1,717,-3,850,10,"FL","N978AT",850,"LGA","MKE",134,738,7,17
"76",2013,1,1,719,-2,1017,5,"B6","N562JB",987,"JFK","MCO",147,944,7,19
"77",2013,1,1,723,-2,1013,-4,"UA","N514UA",962,"EWR","PBI",153,1023,7,23
"78",2013,1,1,724,-6,1111,31,"AA","N541AA",715,"LGA","DFW",254,1389,7,24
"79",2013,1,1,724,-1,1020,-10,"AS","N594AS",11,"EWR","SEA",338,2402,7,24
"80",2013,1,1,725,-5,1052,12,"AA","N4WRAA",2083,"EWR","DFW",238,1372,7,25
"81",2013,1,1,727,-3,959,7,"UA","N37462",1162,"EWR","DEN",254,1605,7,27
"82",2013,1,1,728,-4,1041,3,"UA","N488UA",473,"LGA","IAH",238,1416,7,28
"83",2013,1,1,729,-1,1049,-26,"VX","N635VA",11,"JFK","SFO",356,2586,7,29
"84",2013,1,1,732,-3,857,-1,"B6","N304JB",20,"JFK","ROC",64,264,7,32
"85",2013,1,1,732,3,1041,2,"B6","N563JB",1601,"LGA","RSW",167,1080,7,32
"86",2013,1,1,732,47,1011,30,"UA","N37456",1111,"EWR","MCO",145,937,7,32
"87",2013,1,1,733,-3,854,4,"B6","N552JB",44,"JFK","SYR",54,209,7,33
"88",2013,1,1,734,-3,1047,-26,"B6","N625JB",643,"JFK","SFO",350,2586,7,34
"89",2013,1,1,739,-6,918,-12,"AA","N4WPAA",309,"LGA","ORD",137,733,7,39
"90",2013,1,1,739,0,1104,26,"UA","N37408",1479,"EWR","IAH",249,1400,7,39
"91",2013,1,1,741,-4,1038,2,"B6","N633JB",983,"LGA","TPA",158,1010,7,41
"92",2013,1,1,743,13,1107,7,"AA","N338AA",33,"JFK","LAX",358,2475,7,43
"93",2013,1,1,743,-6,1043,-11,"B6","N624JB",341,"JFK","SRQ",164,1041,7,43
"94",2013,1,1,743,13,1059,3,"DL","N3760C",495,"JFK","SEA",349,2422,7,43
"95",2013,1,1,745,0,1135,10,"AA","N336AA",59,"JFK","SFO",378,2586,7,45
"96",2013,1,1,746,0,1119,-10,"UA","N24224",1668,"EWR","SFO",373,2565,7,46
"97",2013,1,1,749,39,939,49,"MQ","N508MQ",3737,"EWR","ORD",148,719,7,49
"98",2013,1,1,752,-3,1041,-18,"DL","N325US",2263,"LGA","MCO",140,950,7,52
"99",2013,1,1,752,2,1025,-4,"UA","N511UA",477,"LGA","DEN",249,1620,7,52
"100",2013,1,1,752,-7,955,-4,"US","N543UW",1733,"LGA","CLT",96,544,7,52
"101",2013,1,1,753,-2,1056,-14,"AA","N3HMAA",2267,"LGA","MIA",157,1096,7,53
"102",2013,1,1,754,-5,1039,-2,"DL","N935DL",2047,"LGA","ATL",126,762,7,54
"103",2013,1,1,754,-1,1103,33,"WN","N789SW",733,"LGA","DEN",279,1620,7,54
"104",2013,1,1,758,-2,1053,-1,"B6","N645JB",517,"EWR","MCO",142,937,7,58
"105",2013,1,1,759,-1,1057,-30,"DL","N955DL",1843,"JFK","MIA",158,1089,7,59
"106",2013,1,1,800,0,1022,8,"DL","N317US",2119,"LGA","MSP",171,1020,8,0
"107",2013,1,1,800,-10,949,-6,"MQ","N828MQ",4406,"JFK","RDU",80,427,8,0
"108",2013,1,1,801,-4,900,-19,"B6","N206JB",1172,"EWR","BOS",38,200,8,1
"109",2013,1,1,803,-7,903,-22,"AA","N3GEAA",1838,"JFK","BOS",38,187,8,3
"110",2013,1,1,803,3,1132,-12,"UA","N510UA",223,"JFK","SFO",369,2586,8,3
"111",2013,1,1,804,-6,1103,-13,"DL","N947DL",1959,"JFK","MCO",147,944,8,4
相关推荐
- Docker 47 个常见故障的原因和解决方法
-
【作者】曹如熙,具有超过十年的互联网运维及五年以上团队管理经验,多年容器云的运维,尤其在Docker和kubernetes领域非常精通。Docker是一种相对使用较简单的容器,我们可以通过以下几种方式...
- 电脑30个快问快答,解决常见电脑问题
-
1.强行关机/停电对电脑有影响吗?答:可能损坏硬盘(机械硬盘风险高)、未保存数据丢失,偶尔一次影响小,但频繁操作会缩短硬件寿命。2.C盘满影响速度吗?答:会!系统运行需C盘空间缓存临时数据,空间不...
- 使用Tcpdump包抓取分析数据包的详细用法
-
TcpDump可以将网络中传送的数据包的“头”完全截获下来提供分析。它支持针对网络层、协议、主机、网络或端口的过滤,并提供and、or、not等逻辑语句来帮助你去掉无用的信息。tcpdump就是一种...
- 电脑启动不了(BootDevice Not Found Hard Disk-3F0)解决方案
-
HP品牌机,开机启动不了,黑屏,开机取下主板电池恢复BIOS后,开机显示找不到启动盘。一、按F2键进入BIOS,出现硬盘内存检测界面的话,直接退出。就会出现这个界面,光标键向下,选择BIOSSetu...
- 电脑开机黑屏别慌!快码住!起底维修老师傅不能说的秘密
-
按下开机键却只收获黑屏大礼包?那些神秘的英文提示、刺耳的蜂鸣声,其实是电脑在给你发送求救信号!从按下电源到进入桌面的12秒里,你的电脑经历了史诗级的硬件自检与系统加载,今天我们就破译这段“摩斯电码”。...
- 电脑启动故障为何总要先看BIOS?新手必读的关键知识解析
-
最近在帮朋友们解答电脑无法正常开机的问题时,发现大家经常收到一句高频建议:“先检查BIOS”。对不少普通用户而言,BIOS依然是个神秘的存在。那么,BIOS到底是什么?电脑出现哪些故障会与它相关呢?本...
- Windows 11 KB5053598更新:安全补丁还是系统噩梦?
-
2025年3月11日,微软发布了Windows1124H2的强制性更新KB5053598,作为“周二补丁日”(PatchTuesday)的一部分。然而,这款本应提升系统安全性的更新却引发了广泛的...
- 飞牛OS入门安装遇到问题,如何解决?
-
之前小编尝试了用旧电脑装飞牛OS安装之前特意查了一些硬件要求飞牛OS目前支持主流的x86架构硬件主机需能连网线飞牛OS暂时不支持只有无线网卡的安装貌似很多小伙伴在一开始安装就卡住了那今天咱们汇总分...
- 几种常见的电脑开机黑屏显示白色英文字母解决方法
-
当电脑开机出现黑屏并显示白色英文字母时,通常表示系统启动过程中遇到了错误。以下是几种常见原因及对应的解决方法,按照排查顺序整理:一、检查外接设备与硬件连接可能原因:外接U盘、移动硬盘等未拔出,或内部硬...
- 电脑启动出现问题,为什么都要先检查BIOS?
-
【ZOL中关村在线原创技巧应用】最近在回答问题的时候,总会发现很多朋友都在问“电脑无法正常开机怎么办?”这样类似的问题,而许多DIY大佬的回复总会出现一条高频建议“先检查BIOS”。但对于许多普通用户...
- 教你怎么用JavaScript检测当前浏览器是无头浏览器
-
什么是无头浏览器(headlessbrowser)?无头浏览器是指可以在图形界面情况下运行的浏览器。我可以通过编程来控制无头浏览器自动执行各种任务,比如做测试,给网页截屏等。为什么叫“无头”浏览器?...
- 12个高效的Python爬虫框架,你用过几个?
-
实现爬虫技术的编程环境有很多种,Java、Python、C++等都可以用来爬虫。但很多人选择Python来写爬虫,为什么呢?因为Python确实很适合做爬虫,丰富的第三方库十分强大,简单几行代码便可实...
- 运维的报表之路,用 node.js 轻松发送 grafana 报表
-
在运维过程中,无论是监控还是报表,都会有一些通过邮件发送图表的需求,由于开源的zabbix,grafana和kibana等并不完全具有“想发送哪儿就发送哪儿”的图片生成功能,在grafana...
- C#基于浏览器内核的高级爬虫(c#爬取网页内容)
-
基于C#.NET+PhantomJS+Sellenium的高级网络爬虫程序。可执行Javascript代码、触发各类事件、操纵页面Dom结构、甚至可以移除不喜欢的CSS样式。很多网站都用Ajax动态加...
- 如何优化一个秒杀项目?(秒杀实现思路)
-
问题1:使用jmeter性能压测,定位瓶颈代码步骤流程:线程组--->Http请求--->查看结果树--->聚合报告tips:host的文件--->优先调用映射,减少DNS的时...
- 一周热门
-
-
Python实现人事自动打卡,再也不会被批评
-
Psutil + Flask + Pyecharts + Bootstrap 开发动态可视化系统监控
-
一个解决支持HTML/CSS/JS网页转PDF(高质量)的终极解决方案
-
【验证码逆向专栏】vaptcha 手势验证码逆向分析
-
再见Swagger UI 国人开源了一款超好用的 API 文档生成框架,真香
-
网页转成pdf文件的经验分享 网页转成pdf文件的经验分享怎么弄
-
C++ std::vector 简介
-
python使用fitz模块提取pdf中的图片
-
《人人译客》如何规划你的移动电商网站(2)
-
Jupyterhub安装教程 jupyter怎么安装包
-
- 最近发表
- 标签列表
-
- python判断字典是否为空 (50)
- crontab每周一执行 (48)
- aes和des区别 (43)
- bash脚本和shell脚本的区别 (35)
- canvas库 (33)
- dataframe筛选满足条件的行 (35)
- gitlab日志 (33)
- lua xpcall (36)
- blob转json (33)
- python判断是否在列表中 (34)
- python html转pdf (36)
- 安装指定版本npm (37)
- idea搜索jar包内容 (33)
- css鼠标悬停出现隐藏的文字 (34)
- linux nacos启动命令 (33)
- gitlab 日志 (36)
- adb pull (37)
- table.render (33)
- uniapp textarea (33)
- python判断元素在不在列表里 (34)
- python 字典删除元素 (34)
- react-admin (33)
- vscode切换git分支 (35)
- vscode美化代码 (33)
- python bytes转16进制 (35)