upstream/mercurial-mirror Files · mercurial/dagutil.py

demandimport: replace more references to _demandmod instances...

demandimport: replace more references to _demandmod instances _demandmod instances may be referenced by multiple importing modules. Before this patch, the _demandmod instance only maintained a reference to its first consumer when using the "from X import Y" syntax. This is because we only created a single _demandmod instance (attached to the parent X module). If multiple modules A and B performed "from X import Y", we'd produce a single _demandmod instance "demandmod" with the following references: X.Y = <demandmod> A.Y = <demandmod> B.Y = <demandmod> The locals from the first consumer (A) would be stored in <demandmod1>. When <demandmod1> was loaded, we'd look at the locals for the first consumer and replace the symbol, if necessary. This resulted in state: X.Y = <module> A.Y = <module> B.Y = <demandmod> B's reference to Y wasn't updated and was still using the proxy object because we just didn't record that B had a reference to <demandmod> that needed updating! With this patch, we add support for tracking which modules in addition to the initial importer have a reference to the _demandmod instance and we replace those references at module load time. In the case of posix.py, this fixes an issue where the "encoding" module was being proxied, resulting in hundreds of thousands of __getattribute__ lookups on the _demandmod instance during dirstate operations on mozilla-central, speeding up execution by many milliseconds. There are likely several other operation that benefit from this change as well. The new mechanism isn't perfect: references in locals (not globals) may likely linger. So, if there is an import inside a function and a symbol from that module is used in a hot loop, we could have unwanted overhead from proxying through _demandmod. Non-global imports are discouraged anyway. So hopefully this isn't a big deal in practice. We could potentially deploy a code checker that bans use of attribute lookups of function-level-imported modules inside loops. This deficiency in theory could be avoided by storing the set of globals and locals dicts to update in the _demandmod instance. However, I tried this and it didn't work. One reason is that some globals are _demandmod instances. We could work around this, but it's a bit more work. There also might be other module import foo at play. The solution as implemented is better than what we had and IMO is good enough for the time being. It's worth noting that this sub-optimal behavior was made worse by the introduction of absolute_import and its recommended "from . import X" syntax for importing modules from the "mercurial" package. If we ever wrote performance tests, measuring the amount of module imports and __getattribute__ proxy calls through _demandmod instances would be something I'd have it check.

Gregory Szorc - - Load All Authors

File last commit:

r25942:015ded09 default


                r26457:7e813050

default

Download file

             dagutil.py
        
                    286 lines
            
             | 8.2 KiB
            
                | text/x-python
            
             |
                PythonLexer
            
             / mercurial / dagutil.py
          
                    History
                
                 |
                  Annotation
                 | Raw
                 |Copy content
                 |Copy permalink

      # dagutil.py - dag utilities for mercurial

      #

      # Copyright 2010 Benoit Boissinot <bboissin@gmail.com>

      # and Peter Arrenbrecht <peter@arrenbrecht.ch>

      #

      # This software may be used and distributed according to the terms of the

      # GNU General Public License version 2 or any later version.

      from __future__ import absolute_import

      from .i18n import _

      from .node import nullrev

      class basedag(object):

          '''generic interface for DAGs

          terms:

          "ix" (short for index) identifies a nodes internally,

          "id" identifies one externally.

          All params are ixs unless explicitly suffixed otherwise.

          Pluralized params are lists or sets.

          '''

          def __init__(self):

              self._inverse = None

          def nodeset(self):

              '''set of all node ixs'''

              raise NotImplementedError

          def heads(self):

              '''list of head ixs'''

              raise NotImplementedError

          def parents(self, ix):

              '''list of parents ixs of ix'''

              raise NotImplementedError

          def inverse(self):

              '''inverse DAG, where parents becomes children, etc.'''

              raise NotImplementedError

          def ancestorset(self, starts, stops=None):

              '''

              set of all ancestors of starts (incl), but stop walk at stops (excl)

              '''

              raise NotImplementedError

          def descendantset(self, starts, stops=None):

              '''

              set of all descendants of starts (incl), but stop walk at stops (excl)

              '''

              return self.inverse().ancestorset(starts, stops)

          def headsetofconnecteds(self, ixs):

              '''

              subset of connected list of ixs so that no node has a descendant in it

              By "connected list" we mean that if an ancestor and a descendant are in

              the list, then so is at least one path connecting them.

              '''

              raise NotImplementedError

          def externalize(self, ix):

              '''return a node id'''

              return self._externalize(ix)

          def externalizeall(self, ixs):

              '''return a list of (or set if given a set) of node ids'''

              ids = self._externalizeall(ixs)

              if isinstance(ixs, set):

                  return set(ids)

              return list(ids)

          def internalize(self, id):

              '''return a node ix'''

              return self._internalize(id)

          def internalizeall(self, ids, filterunknown=False):

              '''return a list of (or set if given a set) of node ixs'''

              ixs = self._internalizeall(ids, filterunknown)

              if isinstance(ids, set):

                  return set(ixs)

              return list(ixs)

      class genericdag(basedag):

          '''generic implementations for DAGs'''

          def ancestorset(self, starts, stops=None):

              if stops:

                  stops = set(stops)

              else:

                  stops = set()

              seen = set()

              pending = list(starts)

              while pending:

                  n = pending.pop()

                  if n not in seen and n not in stops:

                      seen.add(n)

                      pending.extend(self.parents(n))

              return seen

          def headsetofconnecteds(self, ixs):

              hds = set(ixs)

              if not hds:

                  return hds

              for n in ixs:

                  for p in self.parents(n):

                      hds.discard(p)

              assert hds

              return hds

      class revlogbaseddag(basedag):

          '''generic dag interface to a revlog'''

          def __init__(self, revlog, nodeset):

              basedag.__init__(self)

              self._revlog = revlog

              self._heads = None

              self._nodeset = nodeset

          def nodeset(self):

              return self._nodeset

          def heads(self):

              if self._heads is None:

                  self._heads = self._getheads()

              return self._heads

          def _externalize(self, ix):

              return self._revlog.index[ix][7]

          def _externalizeall(self, ixs):

              idx = self._revlog.index

              return [idx[i][7] for i in ixs]

          def _internalize(self, id):

              ix = self._revlog.rev(id)

              if ix == nullrev:

                  raise LookupError(id, self._revlog.indexfile, _('nullid'))

              return ix

          def _internalizeall(self, ids, filterunknown):

              rl = self._revlog

              if filterunknown:

                  return [r for r in map(rl.nodemap.get, ids)

                          if (r is not None

                              and r != nullrev

                              and r not in rl.filteredrevs)]

              return map(self._internalize, ids)

      class revlogdag(revlogbaseddag):

          '''dag interface to a revlog'''

          def __init__(self, revlog):

              revlogbaseddag.__init__(self, revlog, set(revlog))

          def _getheads(self):

              return [r for r in self._revlog.headrevs() if r != nullrev]

          def parents(self, ix):

              rlog = self._revlog

              idx = rlog.index

              revdata = idx[ix]

              prev = revdata[5]

              if prev != nullrev:

                  prev2 = revdata[6]

                  if prev2 == nullrev:

                      return [prev]

                  return [prev, prev2]

              prev2 = revdata[6]

              if prev2 != nullrev:

                  return [prev2]

              return []

          def inverse(self):

              if self._inverse is None:

                  self._inverse = inverserevlogdag(self)

              return self._inverse

          def ancestorset(self, starts, stops=None):

              rlog = self._revlog

              idx = rlog.index

              if stops:

                  stops = set(stops)

              else:

                  stops = set()

              seen = set()

              pending = list(starts)

              while pending:

                  rev = pending.pop()

                  if rev not in seen and rev not in stops:

                      seen.add(rev)

                      revdata = idx[rev]

                      for i in [5, 6]:

                          prev = revdata[i]

                          if prev != nullrev:

                              pending.append(prev)

              return seen

          def headsetofconnecteds(self, ixs):

              if not ixs:

                  return set()

              rlog = self._revlog

              idx = rlog.index

              headrevs = set(ixs)

              for rev in ixs:

                  revdata = idx[rev]

                  for i in [5, 6]:

                      prev = revdata[i]

                      if prev != nullrev:

                          headrevs.discard(prev)

              assert headrevs

              return headrevs

          def linearize(self, ixs):

              '''linearize and topologically sort a list of revisions

              The linearization process tries to create long runs of revs where

              a child rev comes immediately after its first parent. This is done by

              visiting the heads of the given revs in inverse topological order,

              and for each visited rev, visiting its second parent, then its first

              parent, then adding the rev itself to the output list.

              '''

              sorted = []

              visit = list(self.headsetofconnecteds(ixs))

              visit.sort(reverse=True)

              finished = set()

              while visit:

                  cur = visit.pop()

                  if cur < 0:

                      cur = -cur - 1

                      if cur not in finished:

                          sorted.append(cur)

                          finished.add(cur)

                  else:

                      visit.append(-cur - 1)

                      visit += [p for p in self.parents(cur)

                                if p in ixs and p not in finished]

              assert len(sorted) == len(ixs)

              return sorted

      class inverserevlogdag(revlogbaseddag, genericdag):

          '''inverse of an existing revlog dag; see revlogdag.inverse()'''

          def __init__(self, orig):

              revlogbaseddag.__init__(self, orig._revlog, orig._nodeset)

              self._orig = orig

              self._children = {}

              self._roots = []

              self._walkfrom = len(self._revlog) - 1

          def _walkto(self, walkto):

              rev = self._walkfrom

              cs = self._children

              roots = self._roots

              idx = self._revlog.index

              while rev >= walkto:

                  data = idx[rev]

                  isroot = True

                  for prev in [data[5], data[6]]: # parent revs

                      if prev != nullrev:

                          cs.setdefault(prev, []).append(rev)

                          isroot = False

                  if isroot:

                      roots.append(rev)

                  rev -= 1

              self._walkfrom = rev

          def _getheads(self):

              self._walkto(nullrev)

              return self._roots

          def parents(self, ix):

              if ix is None:

                  return []

              if ix <= self._walkfrom:

                  self._walkto(ix)

              return self._children.get(ix, [])

          def inverse(self):

              return self._orig

	Site-wide shortcuts
/	Use quick search box
g h	Goto home page
g g	Goto my private gists page
g G	Goto my public gists page
g 0-9	Goto bookmarked items from 0-9
n r	New repository page
n g	New gist page

	Repositories
g s	Goto summary page
g c	Goto changelog page
g f	Goto files page
g F	Goto files page with file search activated
g p	Goto pull requests page
g o	Goto repository settings
g O	Goto repository access permissions settings
t s	Toggle sidebar on some pages

				# dagutil.py - dag utilities for mercurial
				#
				# Copyright 2010 Benoit Boissinot <bboissin@gmail.com>
				# and Peter Arrenbrecht <peter@arrenbrecht.ch>
				#
				# This software may be used and distributed according to the terms of the
				# GNU General Public License version 2 or any later version.

				from __future__ import absolute_import

				from .i18n import _
				from .node import nullrev

				class basedag(object):
				'''generic interface for DAGs

				terms:
				"ix" (short for index) identifies a nodes internally,
				"id" identifies one externally.

				All params are ixs unless explicitly suffixed otherwise.
				Pluralized params are lists or sets.
				'''

				def __init__(self):
				self._inverse = None

				def nodeset(self):
				'''set of all node ixs'''
				raise NotImplementedError

				def heads(self):
				'''list of head ixs'''
				raise NotImplementedError

				def parents(self, ix):
				'''list of parents ixs of ix'''
				raise NotImplementedError

				def inverse(self):
				'''inverse DAG, where parents becomes children, etc.'''
				raise NotImplementedError

				def ancestorset(self, starts, stops=None):
				'''
				set of all ancestors of starts (incl), but stop walk at stops (excl)
				'''
				raise NotImplementedError

				def descendantset(self, starts, stops=None):
				'''
				set of all descendants of starts (incl), but stop walk at stops (excl)
				'''
				return self.inverse().ancestorset(starts, stops)

				def headsetofconnecteds(self, ixs):
				'''
				subset of connected list of ixs so that no node has a descendant in it

				By "connected list" we mean that if an ancestor and a descendant are in
				the list, then so is at least one path connecting them.
				'''
				raise NotImplementedError

				def externalize(self, ix):
				'''return a node id'''
				return self._externalize(ix)

				def externalizeall(self, ixs):
				'''return a list of (or set if given a set) of node ids'''
				ids = self._externalizeall(ixs)
				if isinstance(ixs, set):
				return set(ids)
				return list(ids)

				def internalize(self, id):
				'''return a node ix'''
				return self._internalize(id)

				def internalizeall(self, ids, filterunknown=False):
				'''return a list of (or set if given a set) of node ixs'''
				ixs = self._internalizeall(ids, filterunknown)
				if isinstance(ids, set):
				return set(ixs)
				return list(ixs)


				class genericdag(basedag):
				'''generic implementations for DAGs'''

				def ancestorset(self, starts, stops=None):
				if stops:
				stops = set(stops)
				else:
				stops = set()
				seen = set()
				pending = list(starts)
				while pending:
				n = pending.pop()
				if n not in seen and n not in stops:
				seen.add(n)
				pending.extend(self.parents(n))
				return seen

				def headsetofconnecteds(self, ixs):
				hds = set(ixs)
				if not hds:
				return hds
				for n in ixs:
				for p in self.parents(n):
				hds.discard(p)
				assert hds
				return hds


				class revlogbaseddag(basedag):
				'''generic dag interface to a revlog'''

				def __init__(self, revlog, nodeset):
				basedag.__init__(self)
				self._revlog = revlog
				self._heads = None
				self._nodeset = nodeset

				def nodeset(self):
				return self._nodeset

				def heads(self):
				if self._heads is None:
				self._heads = self._getheads()
				return self._heads

				def _externalize(self, ix):
				return self._revlog.index[ix][7]
				def _externalizeall(self, ixs):
				idx = self._revlog.index
				return [idx[i][7] for i in ixs]

				def _internalize(self, id):
				ix = self._revlog.rev(id)
				if ix == nullrev:
				raise LookupError(id, self._revlog.indexfile, _('nullid'))
				return ix
				def _internalizeall(self, ids, filterunknown):
				rl = self._revlog
				if filterunknown:
				return [r for r in map(rl.nodemap.get, ids)
				if (r is not None
				and r != nullrev
				and r not in rl.filteredrevs)]
				return map(self._internalize, ids)


				class revlogdag(revlogbaseddag):
				'''dag interface to a revlog'''

				def __init__(self, revlog):
				revlogbaseddag.__init__(self, revlog, set(revlog))

				def _getheads(self):
				return [r for r in self._revlog.headrevs() if r != nullrev]

				def parents(self, ix):
				rlog = self._revlog
				idx = rlog.index
				revdata = idx[ix]
				prev = revdata[5]
				if prev != nullrev:
				prev2 = revdata[6]
				if prev2 == nullrev:
				return [prev]
				return [prev, prev2]
				prev2 = revdata[6]
				if prev2 != nullrev:
				return [prev2]
				return []

				def inverse(self):
				if self._inverse is None:
				self._inverse = inverserevlogdag(self)
				return self._inverse

				def ancestorset(self, starts, stops=None):
				rlog = self._revlog
				idx = rlog.index
				if stops:
				stops = set(stops)
				else:
				stops = set()
				seen = set()
				pending = list(starts)
				while pending:
				rev = pending.pop()
				if rev not in seen and rev not in stops:
				seen.add(rev)
				revdata = idx[rev]
				for i in [5, 6]:
				prev = revdata[i]
				if prev != nullrev:
				pending.append(prev)
				return seen

				def headsetofconnecteds(self, ixs):
				if not ixs:
				return set()
				rlog = self._revlog
				idx = rlog.index
				headrevs = set(ixs)
				for rev in ixs:
				revdata = idx[rev]
				for i in [5, 6]:
				prev = revdata[i]
				if prev != nullrev:
				headrevs.discard(prev)
				assert headrevs
				return headrevs

				def linearize(self, ixs):
				'''linearize and topologically sort a list of revisions

				The linearization process tries to create long runs of revs where
				a child rev comes immediately after its first parent. This is done by
				visiting the heads of the given revs in inverse topological order,
				and for each visited rev, visiting its second parent, then its first
				parent, then adding the rev itself to the output list.
				'''
				sorted = []
				visit = list(self.headsetofconnecteds(ixs))
				visit.sort(reverse=True)
				finished = set()

				while visit:
				cur = visit.pop()
				if cur < 0:
				cur = -cur - 1
				if cur not in finished:
				sorted.append(cur)
				finished.add(cur)
				else:
				visit.append(-cur - 1)
				visit += [p for p in self.parents(cur)
				if p in ixs and p not in finished]
				assert len(sorted) == len(ixs)
				return sorted


				class inverserevlogdag(revlogbaseddag, genericdag):
				'''inverse of an existing revlog dag; see revlogdag.inverse()'''

				def __init__(self, orig):
				revlogbaseddag.__init__(self, orig._revlog, orig._nodeset)
				self._orig = orig
				self._children = {}
				self._roots = []
				self._walkfrom = len(self._revlog) - 1

				def _walkto(self, walkto):
				rev = self._walkfrom
				cs = self._children
				roots = self._roots
				idx = self._revlog.index
				while rev >= walkto:
				data = idx[rev]
				isroot = True
				for prev in [data[5], data[6]]: # parent revs
				if prev != nullrev:
				cs.setdefault(prev, []).append(rev)
				isroot = False
				if isroot:
				roots.append(rev)
				rev -= 1
				self._walkfrom = rev

				def _getheads(self):
				self._walkto(nullrev)
				return self._roots

				def parents(self, ix):
				if ix is None:
				return []
				if ix <= self._walkfrom:
				self._walkto(ix)
				return self._children.get(ix, [])

				def inverse(self):
				return self._orig