Welcome to our community

Be a part of something great, join today!

  • Hey all, just changed over the backend after 15 years I figured time to give it a bit of an update, its probably gonna be a bit weird for most of you and i am sure there is a few bugs to work out but it should kinda work the same as before... hopefully :)

r3datamanger checksums really that important?

I have very little interest in checksums. Here's my reasoning and methodology, and please tell me if I've missed something here.

I copy the Reddrive straight to backup 1. I unplug the Reddrive and then copy backup 1 to backup 2. I then watch back my backup 2. Every second of it. Looking out for those red dropped frames in RedCine. If every second of this plays back ok, then I will erase the mag and send it back to camera, confident that both backups are good. Any problems, I can go back to the RedDrive and establish whether these are problems created through straight dragging and dropping (not happened once in two years) or whether they are corrupt at source.

I insist on 'checking the gate' - playing back at least a few seconds worth of footage - on every good take, and before every Magazine change. I have never had a situation where footage would play back in camera but fail at the computer.

So what could checksums do to improve my methodology? They will tell me that Backup 2 is identical to the footage from the RedDrive, but they can't tell me that the original footage has dropped frames or has other corruption. So they might just be telling me that I have a perfect copy of some corrupt data. I'd rather be able to mention problems as soon as possible - as soon as I've watched that footage back. And I can sleep at night knowing that I've watched every second of footage back and it's good.

If there are too many cameras to allow this to be possible then checksums become much more useful to me, but still, when I go to sleep that night, I can't promise that there isn't corrupt footage. I can only guarantee that I didn't corrupt it. And that's the problem for me.

The time a checksum costs me could be spent watching this footage back with my own eyes. A good relationship with the camera team can make this entirely possible and realistic.

Tom Turley
www.filmtom.com
 
I have very little interest in checksums. Here's my reasoning and methodology, and please tell me if I've missed something here.

Ok - heres what you missed.

You, as the DIT/DMT, have properly done your job for that day. All the data is copied to 2 different filesystems and you watched every take. You cashed your paycheck, and now you are on to the next job. As far as you are concerned, everything went fine.

The editors take your footage from backup2, copy it to their raid systems and edit away. Months later they get a locked edit for that feature, and now go back to the R3Ds for confirming. How do they know the R3Ds are valid copies?

They could use your method and watch each take again - in real time. Or in my case, since I copied the footage with R3D Data Manager and created a checksum for each take, the editors can validate each take as fast as thier raid system will allow - way faster than realtime. The editors know that they have a mathematically exact copy of the original camera footage without spending days having to watch each frame.

Now, years later, the producer decides to release a directors cut. How do they know the footage is valid? With your method they would again have to watch the footage. However, with checksumed copies, they would be able to pull the checksums and quickly know - in an automated fashion, no humans sitting around staring at computer screens - if the copies are valid.

How about an issue where a hard drive dies during editorial? Say the footage from your backup2 is half copied over to the editorial raid system when the backup drive dies. Do we need to re-copy all the footage? With a checksum you would be able to check at any time that the footage is correct to a mathematical certainty. With your system you would have to guess - even if it is an educated guess - as to which files are valid.

These are all real-world situations that have either happened to me or to productions I have setup. And because they all use checksums I not only sleep good tonight, but every night.

In summary, I think this thread shows that as the DIT/DMT we really need to consider the bigger picture. The whole point of checksums is not just to make sure that one copy is valid, but that every copy is valid. As the DIT/DMT you need to remember that the footage exists long after the job ends. In the overwhelming majority of cases the footage will be copied again at some point. I feel that as the DIT/DMT you should adequately prepare the footage for every copy that will ever be made - and the only way to do that is to use checksums.

The copies you make on set as the DIT/DMT are just the first in a very long series of copies that will be made. Therefore, I believe that as the DIT/DMT and as the only one who has access to the original camera negatives, it is your responsibility to ensure to the best of your ability that at any point in the future anyone can easily check to ensure their footage is still valid.

So what could checksums do to improve my methodology? They will tell me that Backup 2 is identical to the footage from the RedDrive, but they can't tell me that the original footage has dropped frames or has other corruption. So they might just be telling me that I have a perfect copy of some corrupt data. I'd rather be able to mention problems as soon as possible - as soon as I've watched that footage back. And I can sleep at night knowing that I've watched every second of footage back and it's good.

...

The time a checksum costs me could be spent watching this footage back with my own eyes. A good relationship with the camera team can make this entirely possible and realistic.

As I have said before in this thread and elsewhere, checksums are a part of the 3 stages of verification for each copy. For each copy you need to verify that all the files were copied, verify that all the data was copied and verify that the data is valid image data. Only with all three steps done can you ensure that your copy is valid.

So you must watch your footage back as part of the verification.

But checksums can actually save time, as the checksums are done at the speed of your system, not realtime. So if you create a checksum you will then know that each copy is to a mathematical certainty exactly the same. Then it doesnt matter which order you copied to or which order you visually verified, as you have mathematically confirmed that the data in each is the same.

In addition, it seems that people assume its an either/or proposition. Either they checksum or the visually verify, but there isnt time for both. To that, I simply point at the hundreds of productions that I have been on and others around the world that are doing both right now. There is sufficient time, given the challenges of on-set data management, with modern and properly setup hardware to always do checksums. It does exist and is not all that burdensome. In addition, the time and peace of mind that it will save as later copies are made of the footage makes up for any additional cost of resources or manpower needed for the checksums. In the long and medium term, it saves money, at the cost of an extra read cycle today.
 
In terms of delivering verified rushes to the post house and sleeping soundly at night then, we are in agreement.

You, as the DIT/DMT, have properly done your job for that day. All the data is copied to 2 different filesystems and you watched every take. You cashed your paycheck, and now you are on to the next job. As far as you are concerned, everything went fine.

Everything did go fine. The rushes were delivered safely, and with the integrity of the image data verified. Checksums would not have made -specifically- Backup1 and Backup2 any safer.

In each of the situations you list above, running checksums for any copies made from the verified Backup2, after the day has wrapped, in the calmer environment of the post-house would serve in exactly the same way and would be no 'worse' than running checksums on set.

The key to this for me is:
Do the checksums really need to be done on set, when there may be pressure to return the drives or cards to the camera team, where there may be pressure to wrap the location and move on, or wrap for the day?

When I have in the past tried to run checksums on set I have found it to slow down the process too much to be practical on almost any set I've ever worked on. I also feel pressure to watch footage back as soon as possible after they have been delivered to me, because every minute later that I report a corrupted clip, the camera may have moved on to another new position, lighting setups may have changed, costumes changed, actors may have been wrapped etc etc - all of which costs time to put back how it was to reshoot a scene.

Now, in my experience, it may be very very rare that this needs to happen - but I believe we are also in the business of minimising these practical problems that will result from data corruption. The quicker I can let people know about problems with the rushes, the quicker that scene can be reshot, and the less it will cost production.

Perhaps we could conclude that Checksums are not NECESSARY on set:- when under time pressure we must first verify that the data is good to begin with.

But the value of checksums thereafter is another question entirely and, not being from a post-production background myself, one that I will not presume to answer.

Best,

Tom Turley
www.filmtom.com
 
In addition, the CRC checks are usually handled by the drives firmware, which varies by drive to drive and manufacturer to manufacturer. One of the biggest reasons the recent seagate 1.5Tb drives were having so many issues was that the firmware thought the drive was having such a high CRC failure rate that the firmware shut it down to prevent further loss, bricking the drive. Turns out it was a firmware bug, not an issue physically with the drive.

But the overriding point is this. You have your red media and you make your two copies. Now you properly take your one copy and store it, and take your second and use it to edit. Time passes. Days, months, years later you now need access to that footage again. How do you know its valid? CRC checks here wont help you one bit. You could spend days or weeks re-watching all the footage again, but that wont tell you if it is an exact copy. I know that the first footage I transferred shot on build 8 is valid to this day, because I have proper checksums not reliant on any hardware.

I know this is an old thread, but Conrad made an interesting point which I haven't seen anywhere else. If the CRC checksum varies by drive manufacturer, then wouldn't it be done differently by the 2 drives (source and destination)? In that case, the checksum for the initial copy would fail. But you describe a scenario where the the initial checksum matches, then years later it doesn't. How is that possible?
 
If the CRC checksum varies by drive manufacturer, then wouldn't it be done differently by the 2 drives (source and destination)? In that case, the checksum for the initial copy would fail.

As I understand it the CRC checksum is used internally on the drive to verify that the data being read from it is intact before passing it on. It is not comparing a copy against the original. That is where MD5 checksums come in.

But you describe a scenario where the the initial checksum matches, then years later it doesn't. How is that possible?

Data on drives can become corrupted, particularly if a drive sits on a shelf unused. If an MD5 checksum is stored alongside the data, then even if multiple generations of copies are made without performing a checksum of each copy, comparing the MD5 checksum of any copy with the MD5 checksum of the original which travels along with the data gives a very high confidence that a copy is identical to the original. As has been said a number of times on this thread, a visual check is also vital too, since there is no point having a perfect copy of data if that data was in some way corrupt in the first place.
 
Back
Top