In Python, reading a large text file line by line is simple:
for line in open('somefile', 'r'): ...
But how to read the binary file and "split" (by the generator) its contents with some given marker, and not with a new line '\ n'?
I want something like this:
content = open('somefile', 'r').read() result = content.split('some_marker')
but, of course, in terms of memory (the file is about 70 GB). Of course, we cannot read the file for each byte (it will be too slow due to the nature of the hard disk).
The length of the "pieces" (data between these markers) may vary, theoretically from 1 byte to megabytes.
So, to give an example to summarize, the data looks like this (the numbers here indicate bytes, the data is in binary format):
12345223-MARKER-3492-MARKER-34834983428623762374632784-MARKER-888-MARKER-...
Is there an easy way to do this (don't read in chunks, split chunks, remember tails, etc.)?
python
Spaceman
source share