Reading "big" csv files #142

nlhnt · 2022-03-15T15:22:27Z

Hello,
I increased the heapsize to 16G for my kernel with krangl, but reading a CSV file which has 7 columns (int64, str, str, str, int64, str, str) and about 4*e6 rows (almost 800M with utf8 encoding) didn't quite work.
I am stuck with heap size error, Python's pandas was able to load it without much trouble. Unfortunately I face a task where I have to iterate through this table row by row, and Python's loops are not quite useful here (I mean they work, but it not takes a couple of hours to work through that).
Julia's DataFrame.jl was able to load this frame into memory as well (it really is not that big, takes around 2G of RAM on a Windows machine).

holgerbrandl · 2022-04-19T16:54:26Z

This is indeed a problem with the underlying current implementation which is known to not scale well.

@nikitinas Did you come up with a more memory-efficient way to read big tables in https://github.com/Kotlin/dataframe

Soldalma · 2022-11-24T03:58:43Z

I have the same problem with a dataframe of 25000 rows and 700 columns. Actually even after I delete most of the rows the problem persists.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Reading "big" csv files #142

Reading "big" csv files #142

nlhnt commented Mar 15, 2022

holgerbrandl commented Apr 19, 2022

Soldalma commented Nov 24, 2022

Reading "big" csv files #142

Reading "big" csv files #142

Comments

nlhnt commented Mar 15, 2022

holgerbrandl commented Apr 19, 2022

Soldalma commented Nov 24, 2022