Skip to main content
QUICK REVIEW

[Paper Review] Unbiasing Truncated Backpropagation Through Time

Corentin Tallec, Yann Ollivier|arXiv (Cornell University)|May 23, 2017
Topic ModelingComputer Science3 references53 citations
TL;DR

ARTBP introduces stochastic, variable-length truncations with compensation factors to provide unbiased gradient estimates in truncated BPTT, preserving online applicability and improving convergence over standard truncated BPTT. On Penn Treebank character-level modelling, ARTBP slightly improves validation/test performance compared to truncated BPTT.

ABSTRACT

Truncated Backpropagation Through Time (truncated BPTT) is a widespread method for learning recurrent computational graphs. Truncated BPTT keeps the computational benefits of Backpropagation Through Time (BPTT) while relieving the need for a complete backtrack through the whole data sequence at every step. However, truncation favors short-term dependencies: the gradient estimate of truncated BPTT is biased, so that it does not benefit from the convergence guarantees from stochastic gradient theory. We introduce Anticipated Reweighted Truncated Backpropagation (ARTBP), an algorithm that keeps the computational benefits of truncated BPTT, while providing unbiasedness. ARTBP works by using variable truncation lengths together with carefully chosen compensation factors in the backpropagation equation. We check the viability of ARTBP on two tasks. First, a simple synthetic task where careful balancing of temporal dependencies at different scales is needed: truncated BPTT displays unreliable performance, and in worst case scenarios, divergence, while ARTBP converges reliably. Second, on Penn Treebank character-level language modelling, ARTBP slightly outperforms truncated BPTT.

Motivation & Objective

  • Motivate the bias issues in truncated BPTT and the need for unbiased gradient estimates in training RNNs.
  • Introduce ARTBP as a method to achieve unbiasedness while retaining the computational benefits of truncation.
  • Derive the reweighting scheme that compensates for stochastic truncation in backpropagation.
  • Provide theoretical guarantees of unbiased gradient estimates under ARTBP.
  • Empirically validate ARTBP on synthetic tasks and Penn Treebank character-level language modelling.

Proposed method

  • Split training sequences into subsequences with variable truncation lengths sampled from a probability distribution.
  • Modify the backpropagation equation with a compensation factor 1/(1 - c_t) to ensure unbiasedness (Equation 11).
  • Prove that the ARTBP gradient estimate is unbiased (Proposition 1, Equations 12-13).
  • Discuss how to choose truncation probabilities c_t to balance memory and variance (Equation 14).
  • Describe online implementation where updates are made after each subsequence (Section 5).
  • Compare ARTBP to truncated BPTT on synthetic tasks and Penn Treebank (Section 6).

Experimental results

Research questions

  • RQ1Can stochastic, variable-length truncation with appropriate compensation yield unbiased gradient estimates for BPTT?
  • RQ2How does ARTBP trade memory usage and gradient variance relative to fixed-length truncated BPTT?
  • RQ3Do ARTBP and truncated BPTT differ in performance on synthetic tasks requiring multi-timescale dependency learning and real-world language modelling?
  • RQ4What practical guidelines (e.g., choice of c_t) optimize ARTBP’s bias-variance tradeoff for online learning?
  • RQ5Is ARTBP applicable online without backtracking across the entire sequence?

Key findings

  • ARTBP provides unbiased gradient estimates despite using variable truncation lengths.
  • In synthetic tests, truncated BPTT can diverge due to gradient bias, while ARTBP converges reliably.
  • On Penn Treebank character-level language modelling, ARTBP slightly outperforms truncated BPTT in validation and test error.
  • ARTBP introduces gradient variance due to stochastic truncations but reduces memory demands, enabling longer effective traces.
  • A fixed memory-equivalent truncation (L) is comparable to ARTBP with c_t chosen to yield similar average subsequence length, yet ARTBP often yields better convergence properties in biased scenarios.

Better researchstarts right now

From reading papers to final review, dramatically reduce your research time.

No credit card · Free plan available

This review was created by AI and reviewed by human editors.