Welcome to our community

Be a part of something great, join today!

  • Hey all, just changed over the backend after 15 years I figured time to give it a bit of an update, its probably gonna be a bit weird for most of you and i am sure there is a few bugs to work out but it should kinda work the same as before... hopefully :)

What CPU/GPU for 8K Helium playback?

AndreasOberg

Well-known member
Joined
Oct 30, 2011
Messages
1,674
Reaction score
31
Points
48
Location
Leicestershire, United Kingdom
Website
www.obergwildlife.com
Hiya,

With our new Helium I cannot playback full debayer in realtime. We have 2x14 2.7GHz (running at 3.1GHz) and 2xTitan X cards.
My first thought would be that the GPUs were the limit, but I can see that the CPUs are running at 100%. It seems as if we would need maybe 30% CPU power to be able to run if I measure how fast it catches. Sometimes it is faster sometimes slower, probably based on the data.
/Andreas
 
RR-X very helpful for decode (not debayer, alas). That will fix your CPU load problem.
 
You're saying you have 28 cores now and it doesn't work for realtime? What architecture? If so maybe a new 18core won't help...
 
Make sure your storage is delivering at least 500MB/sec sustained read speed.
 
Also it really depends on what software your using and how it uses those resources.

This is indeed still one of the biggest problems with current day software.
Davinci Resolve 14.3 was only able to use 16c/32t (don't know if it is improved in 15) at almost 100%, when you have more cores/threads they do near to nothing.
https://forum.blackmagicdesign.com/viewtopic.php?f=21&t=70646
Most modern day workstation software can only use about 6c/12t with some exceptions like Cinema4D. Davinci Resolve is also pretty well multi-threaded compared to most competitors.

With the Der8auer 18 core you have the best of both worlds:
- Tested with a burn-in test.
- 24 months warrenty.
- 18 cores with a 4.5(4.7)GHz clock (many cores and high frequency so it covers all software) which is faster than the fastest XEON.
- 120 GB/s mem bandwith with DDR4-4000 CL19 (same memory bandwith as the XEON 8180 but with less true latency ~ 10 ns. vs. 14 ns.).
- You need less memory than with XEON's for the same overall speed because of the fewer cores (and you are still able to supply enough memory with memory hungry apps like After Effects).
- With the right motherboard you also have a lot op PCIe-lanes https://www.asus.com/uk/Motherboards/WS-X299-SAGE/ Asus WS X299 SAGE.
- It isn't cheap, but a lot cheaper than the XEON way.
 
Ya I see people complain about their setups sometimes and why they are not getting the best performance but they fail to take into effect how the software uses those resources. Throwing the most expensive CPU or GPU doesn't always mean the best results.

Actually the 7960x outperforms the 7980xe in Resolve. Even the 7940x does well because of the extra speed and for the $ it's within a couple % of performance.

It would be interesting to see a test with extreme overclocked memory and how much of a difference if at all it makes.
 
In what software? and in what timeline resolution. Also define playback so we are talking about the same thing=) I have played back a 12 minute 8k clip in adobe premiere with no dropped frames but Im sure I would drop frames if I started to playback forward, backward, jumping around in the timeline and adding cross dissolves. This was done in a 1080 time line with the 8k 6:1 scaled down to 1080 (set to frame size, not the cheating downscale option). I only have one GTX 1080TI 11g card in this workstation.
 
Ya I see people complain about their setups sometimes and why they are not getting the best performance but they fail to take into effect how the software uses those resources. Throwing the most expensive CPU or GPU doesn't always mean the best results.

Actually the 7960x outperforms the 7980xe in Resolve. Even the 7940x does well because of the extra speed and for the $ it's within a couple % of performance.

It would be interesting to see a test with extreme overclocked memory and how much of a difference if at all it makes.

In Resolve 14.3 it does (in line with the puget test).
(turbo clock x all-cores)
7940x 14 cores at 3.8 GHz 14x3.8= 53.2 92.4%
7960x 16 cores at 3.6 GHz 16x3.6= 57.6 100%
7980xe 16 cores at 3.4 GHz (18-2)x3.4 = 54.4 94.4%

Der8auer 7980xe (18-2)x4.7 = 75.2 130.6%

I don't know how many cores Davinci Resolve 15 can use.

Puget uses DDR4-2666 CL19 for their tests (true latency 14.25 ns).
A set of DDR4-3800 CL19 would give a true latency of 10 ns.
 
Sorry this is in Windows using Redcine X. In Davinci 14.3 I get worse results

Hardware is 2x Xeon e5-2697ver3
Raid has about 750MB/s read
Windows 10
128GB ram
Latest Redcine X
1920x1200 resolution monitor

I tested a dual 10 core 3.1 at work with 1x 1080 ti and it was faster which points to our Titans being the weak lin
 
Last edited:
Sorry this is in Windows using Redcine X. In Davinci 14.3 I get worse results

Hardware is 2x Xeon e5-2697ver3
Raid has about 750MB/s read
Windows 10
128GB ram
Latest Redcine X
1920x1200 resolution monitor

I tested a dual 10 core 3.1 at work with 1x 1080 ti and it was faster which points to our Titans being the weak lin

Titan X Computing Power (FP32) 6,144 GFLOPS, Memory Bandwidth 336 GB/s (all at base freq.)
GTX1080ti Computing Power (FP32) 10,609 GFLOPS Memory Bandwidth 484.35 GB/s (all at base freq.)

GTX1080ti is a lot faster and also has a higher memory bandwidth, so it looks like you found your bottle neck.
 
Just a quick update. I have compared at home with Both 1xTitan X, 2xTitan X and 1x1080 Ti and 2x1080 Ti.
The quick result is that the 2x1080 TI actually finnishes buffering a clip in ideantical speed to 2xTitan X so in that way it seems like the CPU is the limit.

So first conclusion would be that there is no difference, however the reality is a bit different. The 1080 still playbacks it better. The Titan X always stutters a bit even when the material has been buffered, bt with 2x1080 Ti it plays back almost perfectly. With 1x1080 Ti buffering is a bit slower and playback stutters a bit, because the GPU is now the limit.

Andreas
 
Hiya,
More detailed benchmarks:
41533232714_0cbef223d5_b.jpg


How is the Rocket X? I find it a bit frustrating to work with 8k videos right now from a workflow perspective.
 
The original Titan X cards are only 45% as powerful as a 1080TI. And Xeon E5v3 CPUs are crippled by memory speed. And depending on your motherboard and memory configuration, you may even be operating at 1666MHz for memory clock rather than the more ideal 1866Mhz, which is still slow compared to current DDR4 at 2100~2400MHz.

Unfortunately I don’t know if there’s a whole lot you can do here. Updating to newer GPUs will help to an extent, but Redcine-X does a crap job using more than one GPU. And typically the GPU is not the bottleneck in RC-X since it’s very poorly multithreaded as it’s better optimized for the Rocket-X. Your Resolve performance seems a bit slow and is most likely a memory bottleneck. My primary Resolve system has been a dual 10-core E5v2 (20 cores at 3.1GHz) for several years now. I upgraded GPUs to GTX1080s about a year ago and that helped breathe some more life into the system. However, this system chokes on 8K R3Ds... 6K is like butter. I will upgrade with a new system probably next year when Intel refreshes the Xeon platform. Or I may just go for a new box built on a 9th-gen i7 with a bunch of cores and save me a ton of money. My i7 workstations are killing it lately, same with the new iMac Pro systems we have here. Not feeling the whole dual-CPU vibe these days. Not worth the extra expense when the software won’t step up to truly use it. Made lots of sense when it took dual CPUs to get to 16~20 cores. Not so much now... Only makes sense for specific solutions, mostly rendering or complex simulations where it’s easier to distribute load out to all those cores and there are software tools designed to make that happen.
 
How is the Rocket X? I find it a bit frustrating to work with 8k videos right now from a workflow perspective.

Personally, I’m not a fan of the Rocket-X. It’s starting to feel dated compared to the latest GPU hardware and while it does have the edge, there is too much it does not do. Lots of alternate processing functions and R3D extensions have come along that the Rocket-X won’t accelerate. Some people still swear by them though and it’s not to hard to pick one up used for a reasonable price. I would definitely be using a couple Rocket-X cards if I were trying to keep up with rushes on set whatnot. But with Weapon cameras recording DNX or ProRes along with the R3D this has become a non issue as well.

...On that note, I’m finding that while I was an all native workflow at 6K and lower R3Ds, I have to run much lower debayer settings for 8K. Still not a problem to spot check at full debayer as I go and it’s not like I have an 8K monitor anyway. My finishes are still all 4K.... All things considered, it’s not a big deal. But yeah, I’m still a pixel-peeping-geek who would really like to work with and monitor 8K in all its glory. Someday soon, hopefully.
 
Hi Jeff,
Are you not mixing up the Titan X with the old Titan? The Titan X is closer to the 1080 Ti.
Yes I agree it seems the CPU landscape is changing a bit. I especially like the high frequencies with the new i9 CPUs.
Davinci playback I think there is something that is not so optimized with Davinci, it is not using much of the GPU or CPU when in 8k which doesn't make any sense. Some on the forums says that 14.1 was much faster for Helium and after that it got slower.

We are going to build a second edit station and my plan is to go for an i9. I was thinking 14 core before, but after looking at some benchmarks maybe 18 core would be better and then couple it with a Titan 1080 Ti or maybe next generation of card.
/Andreas
 
Last edited:
Also it really depends on what software your using and how it uses those resources.

^ This

I'm no expert but certain softwares are only written to utilize a certain number of cores, so it actually benefits you to have higher clock speeds rather than more cores.
 
Hi Jeff,
Are you not mixing up the Titan X with the old Titan? The Titan X is closer to the 1080 Ti.
Yes I agree it seems the CPU landscape is changing a bit. I especially like the high frequencies with the new i9 CPUs.
Davinci playback I think there is something that is not so optimized with Davinci, it is not using much of the GPU or CPU when in 8k which doesn't make any sense. Some on the forums says that 14.1 was much faster for Helium and after that it got slower.

We are going to build a second edit station and my plan is to go for an i9. I was thinking 14 core before, but after looking at some benchmarks maybe 18 core would be better and then couple it with a Titan 1080 Ti or maybe next generation of card.
/Andreas

The answer was a computex AMD Threadripper 2 32 core

RCX, Davinci Resolve and Cinema4d are compareble in their multi thread performance so cinebenchR15 (MT) should give you a good indication of the performance you can expect.

Stock speeds of current CPU

Intel i7 8700k 95W TDP 1447 points
AMD R7 2700 65W TDP 1529 points
AMD R7 2700x 105W TDP 1817 points
AMD TR 1950x 180W TDP 2986 points
Intel i9 7960x 165W TDP 3161 points
Intel i9 7980xe 165W TDP 3455 points

AMD TR2 32 core 240W TDP 6560 points (est. derived from R7 2700x info times 4) at 3.6 GHz with SMT enabled 32 cores/64 threads
AMD TR2 32 core 240W TDP 4840 points (est. derived from R7 2700x info times 4) at 3.6 GHz with SMT disabled 32 cores/32 threads when software doesn't scale over more the 32 threads (Davinci 14.3)
AMD TR2 32 core 410W TDP 7260 points (est. derived from R7 2700x info times 4) at 4.0 GHz with SMT enabled 32 cores/64 threads
AMD TR2 32 core 410W TDP 5376 points (est. derived from R7 2700x info times 4) at 4.0 GHz with SMT disabled 32 cores/32 threads when software doesn't scale over more the 32 threads (Davinci 14.3)

60+4 PCIe 3 for 3 GPU's, 3 NVMe's, 1 Decklink card.

In other words, I would wait for the release of the TR 2 in august..september 2018 when you want the best bang for the buck
 
Misha I am looking forward to seeing real world benchmarks from the 32 core TR2 but 6560 cinebench? Doubt it. The current 32 core EPYC chips don't come close to that.
 
Back
Top