How to read a large binary file and break its contents into a marker

In Python, reading a large text file line by line is simple:

for line in open('somefile', 'r'): ... 

But how to read the binary file and "split" (by the generator) its contents with some given marker, and not with a new line '\ n'?

I want something like this:

 content = open('somefile', 'r').read() result = content.split('some_marker') 

but, of course, in terms of memory (the file is about 70 GB). Of course, we cannot read the file for each byte (it will be too slow due to the nature of the hard disk).

The length of the "pieces" (data between these markers) may vary, theoretically from 1 byte to megabytes.

So, to give an example to summarize, the data looks like this (the numbers here indicate bytes, the data is in binary format):

 12345223-MARKER-3492-MARKER-34834983428623762374632784-MARKER-888-MARKER-... 

Is there an easy way to do this (don't read in chunks, split chunks, remember tails, etc.)?

+7
python
source share
4 answers

There is no magic in Python that does this for you, but it's not hard to write. For example:

 def split_file(fp, marker): BLOCKSIZE = 4096 result = [] current = '' for block in iter(lambda: fp.read(BLOCKSIZE), ''): current += block while 1: markerpos = current.find(marker) if markerpos == -1: break result.append(current[:markerpos]) current = current[markerpos + len(marker):] result.append(current) return result 

The memory usage of this function can be further reduced by turning it into a generator, i.e. converting result.append(...) to yield ... This remains as an exercise for the reader.

+5
source share

The general idea is to use mmap , after which you can re.finditer :

 import mmap import re with open('somefile', 'rb') as fin: mf = mmap.mmap(fin.fileno(), 0, access=mmap.ACCESS_READ) markers = re.finditer('(.*?)MARKER', mf) for marker in markers: print marker.group(1) 

I have not tested, but you may need (.*?)(MARKER|$) or the like.

Then, down to the OS, to provide the necessary access to the file.

+2
source share

I don’t think there is a built-in function for this, but you can read-to-chunks nicely with an iterator to prevent memory inefficiencies, like @ user4815162342 suggestion:

 def split_by_marker(f, marker = "-MARKER-", block_size = 4096): current = '' while True: block = f.read(block_size) if not block: # end-of-file yield current return current += block while True: markerpos = current.find(marker) if markerpos < 0: break yield current[:markerpos] current = current[markerpos + len(marker):] 

Thus, you will not save all the results in memory at once, and you can still repeat it like this:

 for line in split_by_marker(open(filename, 'rb')): ... 

Just make sure each "line" does not take up too much memory ...

+1
source share

Readline itself reads in chunks, breaks chunks, remembers tails, etc. So no.

0
source share

All Articles