Paste a log, a config, a spreadsheet column or any block of text and get every IPv4 and IPv6 address in it — each one with the line number it first appeared on, that line in full, how many times it occurs, and the IANA block it belongs to. Nothing valid is thrown away quietly: every candidate this extractor refuses is counted with the reason it was refused.
Paste up to 50,000 lines, or drop a .txt, .log or .csv file onto the box. IPv4, IPv6, ports, CIDR prefixes and zone ids are all read.
Drop a file here, or paste. The file is read inside your browser — it is never sent anywhere.
Quick Answer: how do you extract IP addresses from text?
Paste the text into the box above and press Extract. Every IPv4 and IPv6 address is pulled out with the line number it came from, how often it appears, and the IANA block it sits in, so private, reserved and documentation addresses are separated from real public traffic before you act on the list. Nothing leaves your browser, and any candidate that is refused is counted with its reason rather than dropped in silence.
Jessica Wright
Cybersecurity Threat Researcher
Jessica works on IP reputation, blacklists and bot verification — the part of the job that starts with a pile of log lines and ends with a rule you can defend.
This tool was built around two things the competing extractors leave out. The first is the line number, because a list of addresses with no way back to the request that produced them is half an answer. The second is the leading-zero reading: the figures in that section were measured on Ubuntu 24.04.4 with glibc 2.39, PHP 8.4.21, Python 3.11.15 and Node v22.22.2 across 1,049 inputs and eight implementations, and the false-positive counts come from 60,662 lines of real package logs and library source, not from an estimate.
Last reviewed 29 September 2026 · Published 29 September 2026
View all articles by Jessica WrightEvery row is one address. The extractor reports the address exactly as it was written and again in its canonical form, the line number it was first seen on together with that line in full, the line it was last seen on, how many times it occurs, whether it is IPv4 or IPv6, any port or prefix length attached to it, and the IANA Special-Purpose block it falls inside with that block's RFC and the registry's own Globally Reachable value.
The line number is the column that changes how the list is used. When you extract IPs from a log file the addresses are only the start of the question; the useful part is the request, the status code or the rule that sat on the same line. Most extractors hand back a bare list and leave you to search the original file again. The most capable of them is asked this in its own FAQ — can I preserve the original context, the line number and surrounding text, of each extracted address? — and answers only that its output is a clean deduplicated list. Checked again on 29 September 2026. Here the line comes with the address, so a suspicious entry is one glance from its context.
| Column | What it means | Why it is there |
|---|---|---|
| Address | The canonical form. IPv6 is normalised to RFC 5952 spelling; a leading-zero IPv4 form is shown with both readings | Two spellings of one address become one row instead of two |
| As found | The exact characters in your text, including brackets, a port or a zone id | So you can find the string again in the source |
| Line | First line number, that line's text, and the last line the address appears on | Takes you back to the request, not just the address |
| Seen | How many times the address occurs in the scanned lines | One hit is a visitor; four hundred is a pattern |
| IANA block | The Special-Purpose block, its name and RFC, e.g. 192.168.0.0/16 Private-Use, RFC 1918 | Separates real public traffic from address space that cannot reach you |
| Reachable | The registry's own Globally Reachable value: Yes, No or N/A | Blocking something that was never globally reachable achieves nothing |
| Notes | Flags: leading zero, version context, CIDR, port, zone id, embedded IPv4, padded hextets | Marks the rows worth a second look before you act on them |
Sorting is deliberately limited to three orders — first seen, most frequent and numeric. If the list itself is the job, and you want it collapsed into ranges, deduplicated across several pasted files or ordered properly across mixed IPv4 and IPv6, that is a different tool: the IP sorter handles ordering and deduplication of an address list. And if the reason you are extracting addresses in the first place is to stop something, the full picture of where a block can be applied is in the guide to blocking an IP address at every layer.
Key fact: the scan cap is 50,000 lines. Past that the extractor reads the first 50,000 and says how many lines it skipped, because a tool that quietly stops reading is worse than one that admits where it stopped.
An IP address extractor is only as good as the spellings it recognises. Real text is messy: an address may be bracketed with a port, carry a zone id, sit inside a CIDR prefix, be written with padded hextets, or be an IPv4 address hiding inside an IPv6 one. All of the following are read, and each is normalised before counting.
| Written as | Form | Read as |
|---|---|---|
66.249.66.1 | IPv4 dotted quad | 66.249.66.1 |
10.0.3.14:8443 | IPv4 with a port | 10.0.3.14, port 8443 |
192.168.0.0/16 | IPv4 CIDR | 192.168.0.0, prefix 16 |
010.1.1.1 | IPv4 with a leading zero | 10.1.1.1 — and 8.1.1.1 flagged as the octal reading |
2001:db8::dead:beef | Compressed IPv6 | unchanged |
2001:0db8:0000:0000:0000:0000:0000:0001 | Fully written IPv6 | 2001:db8::1 |
2001:DB8::1 | Uppercase hex | 2001:db8::1, flagged |
[2001:db8::1]:443 | Bracketed IPv6 with a port | 2001:db8::1, port 443 |
fe80::1%eth0 | Link-local with a zone id | fe80::1, zone eth0 |
[fe80::1%25eth0] | Zone id percent-encoded for a URI (RFC 6874) | fe80::1, zone eth0 |
2001:db8::/32 | IPv6 CIDR | 2001:db8::, prefix 32 |
::ffff:192.0.2.128 | IPv4-mapped IPv6 | unchanged, flagged as embedded IPv4 |
unknown[203.0.113.55] | IPv4 inside square brackets, the way Postfix and Apache write it | 203.0.113.55 |
ip4:192.0.2.0/24 | An SPF mechanism, where the label ends in a digit | 192.0.2.0/24 |
:: and ::1 | Unspecified and loopback | unchanged |
Refusing something quietly is the failure mode that matters, so every refused candidate is tallied with its reason and up to three examples. A run of text only becomes a candidate if it is address-shaped in the first place: three or more numeric groups joined by dots, or two or more colons with hex digits. A word like deadbeef or a lone key:value is never a candidate, so it is never a rejection either.
| Example | Reason reported |
|---|---|
1.27.4 | Fewer than four numeric groups |
1.2.3.4.5 | More than four numeric groups |
10.0.19045.5371 | A group is larger than 255 — this one is a Windows build number |
12:35:02 and 2026:12:34:56 | Looks like a clock time, not an address |
2001:::1 | Three colons in a row |
2001:db8:1:2:3:4:5:6:7 | More than eight groups |
::ffff:192.0.2.999 | The dotted tail is not a valid IPv4 address |
0xA000102 and 0x0a.1.1.1 | Hexadecimal or packed-integer notation — detected and reported, deliberately not decoded |
192.0.2.0/33 | The prefix length is out of range |
00:1a:2b:3c:4d:5e | Looks like a MAC address, not an IPv6 address |
2001:db8:85a3::1:51888 | Looks like IPv6 with a port but no brackets — RFC 3986 needs [address]:port. A four-digit tail such as :443 is not reported this way, because 443 is also a legal hex group and the address is real |
A prefix in the list is read as a block, not expanded into the addresses inside it. Going the other way — turning a first and last address into the smallest set of prefixes that covers them — is what the IP range to CIDR converter is for, and it is usually the next step after a bulk IP extractor run on a firewall log.
IPv6 gets the same treatment as IPv4 rather than being an afterthought. If your reason for being here is to extract IPv6 addresses specifically, every form in the first table above is read, normalised per RFC 5952 and classified against the IANA IPv6 registry, including the link-local and unique-local ranges that dominate internal logs.
Write 010.1.1.1 in a firewall rule and two pieces of software will disagree about which host you meant. Strict parsers refuse it outright. The C library that sits under ping, getent and a great deal of command-line tooling reads a leading zero as octal, so 010 is 8 and the address becomes 8.1.1.1. That is not a bug in either one, and it is not our interpretation. RFC 6943 section 3.1.1 records it:
In specifying the inet_addr() API, the Portable Operating System Interface (POSIX) standard [IEEE-1003.1] defines "IPv4 dotted decimal notation" as allowing not only strings of the form "10.0.1.2" but also allowing octal and hexadecimal, and addresses with less than four parts. For example, "10.0.258", "0xA000102", and "012.0x102" all represent the same IPv4 address in standard "IPv4 dotted decimal" notation.
— RFC 6943, Issues in Identifier Comparison for Security Purposes, section 3.1.1 (Thaler, Ed., May 2013)
Because this is the kind of claim that gets repeated without checking, it was measured. On 29 September 2026, 1,049 inputs were run through eight implementations on one machine — Ubuntu 24.04.4 LTS with glibc 2.39-0ubuntu8.7, PHP 8.4.21, Python 3.11.15 and Node v22.22.2 — giving 8,392 observations. The inputs were the first octet written as every value from 0 to 255 in four paddings (7, 07, 007, 0007) plus a curated set of short, hexadecimal and integer forms, including the three the RFC names.
| What was measured | Result |
|---|---|
Python ipaddress, Python inet_pton, PHP inet_pton, PHP ip2long, PHP filter_var, Node net.isIPv4 | Unanimous on all 1,049 inputs — the strict parsers never disagree with each other |
Python inet_aton against glibc's own resolver path (getent ahostsv4) | Agreed on all 1,049 inputs, so either can stand for "what the C library does" |
| Inputs the strict six refuse and glibc accepts | 539 |
| … of those, inputs where glibc's address differs from the digits read as decimal | 501 |
| The 768 systematically zero-padded first-octet forms | All 768 refused by all six strict parsers |
… of those, refused by glibc as well, because an 8 or a 9 makes the octal invalid (08.1.1.1, 09.1.1.1) | 246 |
… of those, accepted by glibc with a different value than the digits suggest (012.1.1.1 → 10.1.1.1) | 498 |
So when this extractor meets 010.1.1.1, it does not pick a side. It reports 10.1.1.1 as the plain reading, flags the row, and shows 8.1.1.1 as the octal reading. Where the padding cannot be octal at all — 08 or 09 — the row is flagged octal invalid, which is the honest answer: the C library refuses that string entirely, so no rule written with it will ever match anything.
Why it matters: a blocklist line with a padded octet is a rule aimed at a different host than the one you read on the page, or at no host at all. That is a silent failure in exactly the place you cannot afford one.
The same section of the RFC is why hexadecimal and packed-integer forms are reported but not decoded here. 0xA000102 and 012.0x102 both mean 10.0.1.2 to inet_aton — the measurement above confirms it on this machine — but silently converting them would put an address in your list that never appeared in your text. They are counted as rejections with the reason stated instead.
A list of addresses with no idea which ones can reach you is a list you cannot act on. Every address found here is matched, longest prefix first, against the IANA IPv4 Special-Purpose Address Space registry and its IPv6 counterpart. Both were read on 25 September 2026 and each carries its own Last Updated 2025-10-09. The registry's Globally Reachable column is copied as it stands, not derived from anything on this page, and it means:
A boolean value indicating whether an IP datagram whose destination address is drawn from the allocated special-purpose address block is forwardable beyond a specified administrative domain.
— the definition the registry uses, set out as the Global attribute in RFC 6890 section 2.2.1 (April 2013).
These are the blocks that turn up most often in a real log, and what each one means when it appears:
| Block | Name, RFC | Reachable | What it means in your log |
|---|---|---|---|
10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16 | Private-Use, RFC 1918 | No | Internal traffic, or a proxy passing on a client's own LAN address |
100.64.0.0/10 | Shared Address Space, RFC 6598 | No | Carrier-grade NAT. Thousands of subscribers can sit behind one public address |
127.0.0.0/8 | Loopback, RFC 1122 | No | The machine itself — health checks, local scripts, a reverse proxy on the same host |
169.254.0.0/16 | Link Local, RFC 3927 | No | A host that never got a DHCP lease |
192.0.2.0/24, 198.51.100.0/24, 203.0.113.0/24 | Documentation, RFC 5737 | No | Sample output, a copied config or a test fixture, not live traffic |
fe80::/10 | Link-Local Unicast, RFC 4291 | No | IPv6 on the local segment; the zone id tells you which interface |
fc00::/7 | Unique-Local, RFC 4193 | No | The IPv6 equivalent of RFC 1918 space |
2001:db8::/32 | Documentation, RFC 3849 | No | Examples and copied snippets |
3fff::/20 | Documentation, RFC 9637 | No | The newer documentation block, which most tooling has not caught up with |
2000::/3 | Global Unicast, RFC 4291 | Yes | Ordinary public IPv6 |
Anything that matches no special-purpose block is ordinary public unicast, and those are the rows that matter when you are deciding what to keep out. The next question is usually whose network it is: a datacentre range behaves nothing like a residential one, and the trade-offs of drawing that line are covered in blocking datacentre and hosting IPs without breaking your own integrations.
Two results say N/A rather than Yes or No, and the reason is the same in both cases: there is no value to copy.
224.0.0.0/4, ff00::/8) is not in the special-purpose registry at all — it has its own, RFC 5771 — so there is no Globally Reachable column for it. It is also never a source address you would block, so it stays out of the blocklist output.Limit: the registries say what an address is, not who runs it. There is no geolocation, no ASN and no ownership here, because none of that can be answered without a network call — and this tool makes none.
Four numbers under 256 joined by dots is a version number as often as it is an address. Anyone who tries to find IP addresses in text hits this immediately, and the usual response is a confident sentence claiming false positives are rare. Here is a number instead.
A candidate glued to the end of a word is not a candidate. 0ubuntu0.24.04.4 contains 0.24.04.4, which is four groups all under 256 — but it is preceded by a letter, so it is skipped. The single exception is a v prefix standing on its own: v1.2.3.4 really might be an address written carelessly, so it is extracted and flagged version context rather than dropped. That is the whole philosophy of this extractor — a doubtful row you can see beats a real address you never knew was there.
Five corpora, all real files on one machine, chosen because they are known to contain no IP addresses at all. Every acceptance is therefore a false positive by definition. The per-corpus unique counts overlap — the package log repeats the same seven version numbers as the version list — so the total counts distinct strings, not the sum of the column.
| Corpus | Lines | Accepted | Unique | What they were |
|---|---|---|---|---|
Every package version dpkg-query reports on Ubuntu 24.04.4 | 1,011 | 9 | 7 | Four-part version numbers such as 2.3.3.4 and 0.99.49.4 |
Python package versions from pip list | 158 | 3 | 1 | 4.13.0.92 |
| Python 3.11 standard library source, first 50,000 lines | 50,000 | 6 | 4 | :: three times as a reStructuredText marker, two HTML5 spec section numbers, and one address that really is an address inside a docstring |
/var/log/dpkg.log | 9,263 | 63 | 7 | The same four-part version numbers, repeated once per package state change |
/var/log/alternatives.log | 230 | 0 | 0 | Nothing |
| Total | 60,662 | 81 | 12 distinct | 8 four-part version numbers, 2 HTML5 spec section numbers, one reStructuredText ::, one genuine address in a docstring |
Those same corpora produced 15,815 refused candidates, almost all of them either clock times (9,501) or numbers with fewer than four groups (6,160) — which is the other half of the picture: the reject counter tells you what the extractor decided against, so you can judge whether it decided correctly.
Reading the flags: if a row is flagged version context, or its line text is a package list rather than a request, it is almost certainly a version number. The line column is what makes that judgement possible in one glance.
You do not need this page to get addresses out of a file, and it is worth knowing exactly what the command-line answer gives you and where it stops. Every command below was run against the same six-line test log before it was published here, and the output shown is the output it produced.
# -o prints only the match, -n prefixes the line number
grep -Eon '([0-9]{1,3}\.){3}[0-9]{1,3}' access.log
# 1:66.249.66.1
# 3:203.0.113.77
# 4:999.999.999.999 <- not an address
# 5:010.1.1.1
# 5:10.0.3.14
It works, and it is the right reach for a quick look. It also accepts 999.999.999.999, because [0-9]{1,3} says nothing about the range of an octet.
grep -Eon '(25[0-5]|2[0-4][0-9]|1[0-9]{2}|[1-9]?[0-9])(\.(25[0-5]|2[0-4][0-9]|1[0-9]{2}|[1-9]?[0-9])){3}' access.log
# 1:66.249.66.1
# 3:203.0.113.77
# 5:10.1.1.1 <- the input was 010.1.1.1
# 5:10.0.3.14
999.999.999.999 is gone, which is the improvement people reach for. Look at line 5. The file says 010.1.1.1; the regex reports 10.1.1.1, because it simply starts matching after the zero. The leading zero has not been handled, it has been made invisible — and the address the C library would resolve is 8.1.1.1, a third value that appears nowhere in either output.
# occurrence counts, but the line numbers are gone
grep -Eoh '...' access.log | sort | uniq -c | sort -rn
# 2 66.249.66.1
# 1 203.0.113.77
# an IPv6 attempt, on the same file
grep -Eoh '([0-9a-fA-F]{0,4}:){2,7}[0-9a-fA-F]{0,4}' access.log | sort -u
# 2001:db8::dead:beef
# 2026:12:34:56 <- a time stamp
# 2026:12:34:58 <- a time stamp
# 2026:12:35:01 <- a time stamp
Four matches, one address. Nothing shorter than a real parser separates an IPv6 address from an Apache time stamp, because they are the same shape.
import re, ipaddress
pat = re.compile(r'[0-9a-fA-F:.]{2,45}')
seen = {}
for n, line in enumerate(open('access.log'), 1):
for m in pat.finditer(line):
try: a = ipaddress.ip_address(m.group())
except ValueError: continue
seen.setdefault(str(a), n)
# 1 66.249.66.1
# 3 203.0.113.77
# 6 2001:db8::dead:beef
No time stamps, no 999.999.999.999, and a line number. It still misses 10.0.3.14, because the candidate it cut out was 10.0.3.14:8443 and that is not an address to ipaddress; it drops 010.1.1.1 without a word, because the strict parser refuses it; and it will not keep a CIDR prefix, a zone id or a port.
The honest summary: grep is right for a glance, and the Python is right if you are writing a pipeline. The extractor above exists for the gap between them — ports, prefixes, zone ids, both readings of a padded octet, the IANA block, the count and the line, in one pass and without installing anything.
The Blocklist ready output exists because a raw extraction is not a blocklist. It keeps only the addresses the IANA registries mark globally reachable, drops ports and zone ids, keeps any CIDR prefix intact and sorts numerically. Loopback, RFC 1918 space, carrier-grade NAT and documentation ranges are left out on purpose: a firewall rule against 127.0.0.1 or 192.0.2.5 costs you a line of config and blocks nothing.
Three things are worth doing before any of it reaches a live rule:
For reference, the same list in three common formats:
# Apache 2.4 - .htaccess
<RequireAll>
Require all granted
Require not ip 203.0.113.77
Require not ip 198.51.100.0/24
</RequireAll>
# nftables
nft add rule inet filter input ip saddr { 203.0.113.77, 198.51.100.0/24 } drop
# ufw, one address at a time
sudo ufw deny from 203.0.113.77
The CSV output is the one to keep if the list is going into a ticket or a review: it carries the line number, the line text, the occurrence count and the IANA classification, which is the evidence for the rule rather than just the rule.
Being clear about the edges is part of being useful. This tool will not:
10.0.0.0/8 is reported as one block, not 16,777,216 addresses.One honest limit on the scan: above 50,000 lines the extractor stops and reports the shortfall. It does not stream, chunk or sample, because a partial result presented as a complete one is the worst outcome available.
Open the log and paste it into the box above, or drop the .log, .txt or .csv file straight onto the drop area. The extractor reads the text line by line, so an Apache combined log, an nginx error log, a syslog file, a firewall export and a CSV all work without any setting change. Every address comes back with the line number it was first seen on and that line in full, so you can go back to the source and see the request that produced it.
Yes. It reads compressed IPv6 (2001:db8::1), the fully written 39-character form, uppercase hex, a bracketed address with a port ([2001:db8::1]:443), a zone id (fe80::1%eth0), a prefix (2001:db8::/32) and the IPv4-mapped form (::ffff:192.0.2.128). Every one is normalised to the RFC 5952 canonical spelling before it is counted, so two different spellings of the same address collapse into one row instead of appearing twice.
Because two families of software read it two different ways. Strict parsers refuse a leading zero outright; the C library behind ping and much command-line tooling reads a leading zero as octal, so 010 becomes 8. Both readings are shown, and neither is guessed: RFC 6943 section 3.1.1 records the behaviour and the numbers in the leading-zero section of this page were measured on this machine, not assumed. In a blocklist, that difference is the gap between the rule you wrote and the host it hits.
Fifty thousand. Beyond that the extractor scans the first 50,000 lines and tells you plainly how many lines it did not read, rather than truncating in silence or freezing the tab. Fifty thousand lines of a combined access log is a little over 9 MB — a mean of 191 bytes a line, measured over 50,000 generated combined-log lines — and scans in well under a second on a normal laptop. If your file is larger, split it and run it in passes, or filter it first with grep so the part you care about fits.
No. The whole extractor is JavaScript running inside your own browser tab: no network request is made when you press Extract, and dropping a file reads it locally with the browser file API. Nothing is sent to TrustMyIP, nothing is stored, and closing the tab discards it. That is also why the tool works on an internal log you would never be allowed to paste into a hosted service.
Because 1.2.3.4 as a version and 1.2.3.4 as an address are the same four numbers. Anything glued to the end of a word is ignored, so 0ubuntu0.24.04.4 is left alone, and a v prefix is flagged as version context - but a four-group number standing on its own is extracted, because dropping a real address on a guess is worse than showing a doubtful one. The measured rate over real corpora is published on this page, and the flag column tells you which rows to look at twice.
Yes. The CSV output carries every column the table shows - normalised address, the text as found, family, prefix length, zone id, ports, occurrence count, first and last line number, IANA block, block name, RFC, the registry's Globally Reachable value, the octal reading where one exists, the flags and the first line's full text. Cells are quoted and embedded quotes doubled, so a log line containing a comma or a quotation mark opens correctly in Excel, Sheets and LibreOffice.
Because 203.0.113.0/24 is TEST-NET-3, one of three blocks RFC 5737 reserves for documentation, and the IANA registry marks it not globally reachable. Seeing it in a real log usually means you are looking at sample output, a copied config or a test fixture rather than live traffic. The same column is what separates a genuine visitor from loopback, link-local, carrier-grade NAT and RFC 1918 private space before you write a single firewall rule.
The usual order: pull the addresses out, check what they are, collapse them into ranges, then write the rule.
A user agent string proves nothing. Forward and reverse DNS is the only check that separates a genuine crawler from something wearing its name — and it is worth running before a single address from your list reaches a firewall rule.
Last updated 29 September 2026 · classification from the IANA IPv4 and IPv6 Special-Purpose Address registries, read 25 September 2026