Check that you have installed one module, or two modules, one per each memory channel. If two modules are installed in wrong slots, they can use the same channel and this can result in a slowdown.
One a 13900 core, with default clock, should be able to read at max. rate 89.2/2 = 44.8 GB/s (small gigabytes = 10^9). One DDR5 module running at 6000 MT/s has maximum throughput 48 GB/s. Granting some unknown slowdowns due to error corrections and other things I would be expecting to see at least above 40 GB/s. With two RAM modules and two or more cores working, one should be able to get above 80GB/s.
Perhaps memtest does something that limits its memory performance (maybe it uses old less-efficient instructions for reading memory instead of SSE/AVX) or maybe it makes a factor 2 error in calculation of the speed.
I have two slots installed correctly. The bandwidth tool in AUR reports ~20-35 GB/s for sequential reads (although weird spikes of 70 gb/s) i.e. which is consistent with memtest86's ~25 GB/s (the discrepancy can be explained away by memtest probably not using AVX instructions). Writes seem a lot better at ~58 GB/s bypassing cache.
That's still a large discrepancy though where I'm typically 30% away from the nominal speed.
Make sure all power saving stuff is disabled in BIOS, that XMP is being used, and Linux is not booting with some weird ACPI/APIC flags. Try ganged/unganged mode (DCT) in BIOS.
You have to be careful reading the charts. The author of that tool doesn’t do the best job clarifying that most of that chart is showing L1/2/3 cache speeds. That tool could definitely use some TLC. I don’t see any that get anywhere near that once your out of cache range.
First row for Intel Core i7-930, DDR3 2000 MT/s he should be getting at most 16GB/s for reading from DRAM modules (not cache), but he's actually getting 18.4 GB/s, big 15% more. Maybe more cores were used, then the theoretical limit of that processor is 25.6GB/s, and the result is not that great.
But then for i5-520M, 1066MT/s he should be getting 8.5GB/s per core, but he's getting less, only 7.1GB/s.
Maybe contact the author (his email is in README.txt) and ask for clarifications on these discrepancies and help on your problem.
I think there is a problem somewhere with your system/measurements. With 6000 MT/s, there's no way getting 25GB/s is fine - that's what DDR4 3200MT/s should able to do. You should be getting close below 48 GB/s on one core, and 89.6GB/s multicore. And writes should be slower than reads.
Comments
Check that you have installed one module, or two modules, one per each memory channel. If two modules are installed in wrong slots, they can use the same channel and this can result in a slowdown.
One a 13900 core, with default clock, should be able to read at max. rate 89.2/2 = 44.8 GB/s (small gigabytes = 10^9). One DDR5 module running at 6000 MT/s has maximum throughput 48 GB/s. Granting some unknown slowdowns due to error corrections and other things I would be expecting to see at least above 40 GB/s. With two RAM modules and two or more cores working, one should be able to get above 80GB/s.
Perhaps memtest does something that limits its memory performance (maybe it uses old less-efficient instructions for reading memory instead of SSE/AVX) or maybe it makes a factor 2 error in calculation of the speed.
I have two slots installed correctly. The bandwidth tool in AUR reports ~20-35 GB/s for sequential reads (although weird spikes of 70 gb/s) i.e. which is consistent with memtest86's ~25 GB/s (the discrepancy can be explained away by memtest probably not using AVX instructions). Writes seem a lot better at ~58 GB/s bypassing cache.
That's still a large discrepancy though where I'm typically 30% away from the nominal speed.
That's really weird, the author of the bandwidth tool has line graphs with >100GB/s on weaker hardware.
https://zsmith.co/bandwidth.php
Make sure all power saving stuff is disabled in BIOS, that XMP is being used, and Linux is not booting with some weird ACPI/APIC flags. Try ganged/unganged mode (DCT) in BIOS.
You have to be careful reading the charts. The author of that tool doesn’t do the best job clarifying that most of that chart is showing L1/2/3 cache speeds. That tool could definitely use some TLC. I don’t see any that get anywhere near that once your out of cache range.
Oh, you might be right. Not sure what you mean by TLC. But indeed the charts lack an explanation, so I've looked instead into this table:
https://zsmith.co/bw-table.php
First row for Intel Core i7-930, DDR3 2000 MT/s he should be getting at most 16GB/s for reading from DRAM modules (not cache), but he's actually getting 18.4 GB/s, big 15% more. Maybe more cores were used, then the theoretical limit of that processor is 25.6GB/s, and the result is not that great.
But then for i5-520M, 1066MT/s he should be getting 8.5GB/s per core, but he's getting less, only 7.1GB/s.
Maybe contact the author (his email is in README.txt) and ask for clarifications on these discrepancies and help on your problem.
I think there is a problem somewhere with your system/measurements. With 6000 MT/s, there's no way getting 25GB/s is fine - that's what DDR4 3200MT/s should able to do. You should be getting close below 48 GB/s on one core, and 89.6GB/s multicore. And writes should be slower than reads.
People are getting 95GB/s with AIDA64 on Windows.
https://www.thefpsreview.com/2022/10/20/intel-core-i9-13900k...