블로그
- GPU (Full-Sharded Data Parallel) 이 후에는 communication library (Allreduce)나 nccl로 gradient 싱크로나이즈를 통해 아래와 같이 각 GPU는 같은 gradient를 가지게 됩니다 아래는 GPT-2 GPU 메모리 필용량 FSDP의 경우 아래와 같이 작용합.......
- cuDNN NVIDIA cuDNN The NVIDIA CUDA® Deep Neural Network library (cuDNN) is a GPU-accelerated library of primitives training neural networks and developing software applications rather than spending time on low-level GPU
- Failed to get convolution algorithm. This is probably because cuDNN failed to initialize 32.430793: I tensorflow/stream_executor/platform/default/dso_loader.cc:44] Successfully opened dynamic library nvcuda.dll 2020-03-18 20:32:34.664467: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1618] Found