upstream/mercurial-mirror Files · mercurial/pure/parsers.py

manifest: persist the manifestfulltext cache...

manifest: persist the manifestfulltext cache Reconstructing the manifest from the revlog takes time, so much so that there already is a LRU cache to avoid having to load a manifest multiple times. This patch persists that LRU cache in the .hg/cache directory, so we can re-use this cache across hg commands. Commit benchmark (run on Macos 10.13 on a 2017-model Macbook Pro with Core i7 2.9GHz and flash drive), testing without and with patch run 5 times, baseline is r2a227782e754: * committing to an existing file, against the mozilla-central repository. Baseline real time average 1.9692, with patch 1.3786. A new debugcommand "hg debugmanifestfulltextcache" lets you inspect the cache, clear it, or add specific manifest nodeids to it. When calling repo.updatecaches(), the manifest(s) for the working copy parents are added to the cache. The hg perfmanifest command has an additional --clear-disk switch to clear this cache when testing manifest loading performance. Using this command to test performance on the firefox repository for revision f947d902ed91, whose manifest has a delta chain length of 60540, we see: $ hg perfmanifest --clear-disk ! wall 0.972253 comb 0.970000 user 0.850000 sys 0.120000 (best of 10) $ hg debugmanifestfulltextcache -a `hg log --debug -r | grep manifest | cut -d: -f3` Cache contains 1 manifest entries, in order of most to least recent: id: 0294517df4aad07c70701db43bc7ff24c3ce7dbc, size 25.6 MB Total cache data size 25.6 MB, on-disk 0 bytes $ hg perfmanifest ! wall 0.036748 comb 0.040000 user 0.020000 sys 0.020000 (best of 100) Worst-case scenario: a manifest text loaded from a single delta; in the firefox repository manifest node is the chain base for the manifest attached to revision f947d902ed91. Loading this from a full cache file is just as fast as without the cache; the extra node ids ensure a big full cache: $ for node in 1a1922c14a3e 0294517df4aa; do > hgd debugmanifestfulltextcache -a $node > /dev/null > done $ hgd perfmanifest -m ! wall 0.077513 comb 0.080000 user 0.030000 sys 0.050000 (best of 100) $ hgd perfmanifest -m --clear-disk ! wall 0.078547 comb 0.080000 user 0.070000 sys 0.010000 (best of 100)

Alex Gaynor - - Load All Authors

File last commit:

r34332:53133250 default


                r38803:0a57945a

default

Download file

             parsers.py
        
                    179 lines
            
             | 5.5 KiB
            
                | text/x-python
            
             |
                PythonLexer
            
             / mercurial / pure / parsers.py
          
                    History
                
                 |
                  Source
                 | Raw
                 |Copy content
                 |Copy permalink

        Martin Geisler
    
pure Python implementation of parsers.c

              r7700
            
      # parsers.py - Python implementation of parsers.c

      #

      # Copyright 2009 Matt Mackall <mpm@selenic.com> and others

      #

        Martin Geisler
    
updated license to be explicit about GPL version 2

              r8225
            
      # This software may be used and distributed according to the terms of the

        Matt Mackall
    
Update license to GPLv2+

              r10263
            
      # GNU General Public License version 2 or any later version.

        Martin Geisler
    
pure Python implementation of parsers.c

              r7700
            
        Gregory Szorc
    
parsers: use absolute_import

              r27339
            
      from __future__ import absolute_import

      import struct

      import zlib

        Yuya Nishihara
    
parsers: switch to policy importer...

              r32372
            
      from ..node import nullid

      from .. import pycompat

        Gregory Szorc
    
util: prefer "bytesio" to "stringio"...

              r36976
            
      stringio = pycompat.bytesio

        Martin Geisler
    
pure Python implementation of parsers.c

              r7700
            
        Pulkit Goyal
    
parsers: alias long to int on Python 3

              r31220
            
        Martin Geisler
    
pure Python implementation of parsers.c

              r7700
            
      _pack = struct.pack

      _unpack = struct.unpack

      _compress = zlib.compress

      _decompress = zlib.decompress

        Siddharth Agarwal
    
parsers: inline fields of dirstate values in C version...

              r21809
            
      # Some code below makes tuples directly because it's more convenient. However,

      # code outside this module should always use dirstatetuple.

      def dirstatetuple(*x):

          # x is a tuple

          return x

        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
      indexformatng = ">Qiiiiii20s12x"

      indexfirst = struct.calcsize('Q')

      sizeint = struct.calcsize('i')

      indexsize = struct.calcsize(indexformatng)

      def gettype(q):

          return int(q & 0xFFFF)

        Matt Mackall
    
pure/parsers: fix circular imports, import mercurial modules properly

              r7945
            
        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
      def offset_type(offset, type):

        Martin von Zweigbergk
    
pure: use int instead of long...

              r31529
            
          return int(int(offset) << 16 | type)

        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
      class BaseIndexObject(object):

          def __len__(self):

              return self._lgt + len(self._extra) + 1

          def insert(self, i, tup):

              assert i == -1

              self._extra.append(tup)

        Matt Mackall
    
pure/parsers: fix circular imports, import mercurial modules properly

              r7945
            
        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
          def _fix_index(self, i):

              if not isinstance(i, int):

                  raise TypeError("expecting int indexes")

              if i < 0:

                  i = len(self) + i

              if i < 0 or i >= len(self):

                  raise IndexError

              return i

        Matt Mackall
    
pure/parsers: fix circular imports, import mercurial modules properly

              r7945
            
        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
          def __getitem__(self, i):

              i = self._fix_index(i)

              if i == len(self) - 1:

                  return (0, 0, 0, -1, -1, -1, -1, nullid)

              if i >= self._lgt:

                  return self._extra[i - self._lgt]

              index = self._calculate_index(i)

              r = struct.unpack(indexformatng, self._data[index:index + indexsize])

              if i == 0:

                  e = list(r)

                  type = gettype(e[0])

                  e[0] = offset_type(0, type)

                  return tuple(e)

              return r

      class IndexObject(BaseIndexObject):

          def __init__(self, data):

              assert len(data) % indexsize == 0

              self._data = data

              self._lgt = len(data) // indexsize

              self._extra = []

          def _calculate_index(self, i):

              return i * indexsize

        Matt Mackall
    
revlog: remove lazy index

              r13253
            
        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
          def __delitem__(self, i):

        Alex Gaynor
    
style: always use `x is not None` instead of `not x is None`...

              r34332
            
              if not isinstance(i, slice) or not i.stop == -1 or i.step is not None:

        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
                  raise ValueError("deleting slices only supports a:-1 with step 1")

              i = self._fix_index(i.start)

              if i < self._lgt:

                  self._data = self._data[:i * indexsize]

                  self._lgt = i

                  self._extra = []

              else:

                  self._extra = self._extra[:i - self._lgt]

      class InlinedIndexObject(BaseIndexObject):

          def __init__(self, data, inline=0):

              self._data = data

              self._lgt = self._inline_scan(None)

              self._inline_scan(self._lgt)

              self._extra = []

        Martin Geisler
    
pure Python implementation of parsers.c

              r7700
            
        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
          def _inline_scan(self, lgt):

              off = 0

              if lgt is not None:

                  self._offsets = [0] * lgt

              count = 0

              while off <= len(self._data) - indexsize:

                  s, = struct.unpack('>i',

                      self._data[off + indexfirst:off + sizeint + indexfirst])

                  if lgt is not None:

                      self._offsets[count] = off

                  count += 1

                  off += indexsize + s

              if off != len(self._data):

                  raise ValueError("corrupted data")

              return count

        Augie Fackler
    
pure parsers: properly detect corrupt index files...

              r14421
            
        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
          def __delitem__(self, i):

        Alex Gaynor
    
style: always use `x is not None` instead of `not x is None`...

              r34332
            
              if not isinstance(i, slice) or not i.stop == -1 or i.step is not None:

        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
                  raise ValueError("deleting slices only supports a:-1 with step 1")

              i = self._fix_index(i.start)

              if i < self._lgt:

                  self._offsets = self._offsets[:i]

                  self._lgt = i

                  self._extra = []

              else:

                  self._extra = self._extra[:i - self._lgt]

        Martin Geisler
    
pure Python implementation of parsers.c

              r7700
            
        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
          def _calculate_index(self, i):

              return self._offsets[i]

        Martin Geisler
    
pure Python implementation of parsers.c

              r7700
            
        Maciej Fijalkowski
    
pure: write a really lazy version of pure indexObject...

              r29133
            
      def parse_index2(data, inline):

          if not inline:

              return IndexObject(data), None

          return InlinedIndexObject(data, inline), (0, data)

        Martin Geisler
    
pure Python implementation of parsers.c

              r7700
            
      def parse_dirstate(dmap, copymap, st):

          parents = [st[:20], st[20: 40]]

        Mads Kiilerich
    
fix wording and not-completely-trivial spelling errors and bad docstrings

              r17425
            
          # dereference fields so they will be local in loop

        Matt Mackall
    
pure/parsers: fix circular imports, import mercurial modules properly

              r7945
            
          format = ">cllll"

          e_size = struct.calcsize(format)

        Martin Geisler
    
pure Python implementation of parsers.c

              r7700
            
          pos1 = 40

          l = len(st)

          # the inner loop

          while pos1 < l:

              pos2 = pos1 + e_size

              e = _unpack(">cllll", st[pos1:pos2]) # a literal here is faster

              pos1 = pos2 + e[4]

              f = st[pos2:pos1]

              if '\0' in f:

                  f, c = f.split('\0')

                  copymap[f] = c

              dmap[f] = e[:4]

          return parents

        Siddharth Agarwal
    
dirstate: move pure python dirstate packing to pure/parsers.py

              r18567
            
      def pack_dirstate(dmap, copymap, pl, now):

          now = int(now)

        timeless
    
pycompat: switch to util.stringio for py3 compat

              r28861
            
          cs = stringio()

        Siddharth Agarwal
    
dirstate: move pure python dirstate packing to pure/parsers.py

              r18567
            
          write = cs.write

          write("".join(pl))

          for f, e in dmap.iteritems():

              if e[0] == 'n' and e[3] == now:

                  # The file was last modified "simultaneously" with the current

                  # write to dirstate (i.e. within the same second for file-

                  # systems with a granularity of 1 sec). This commonly happens

                  # for at least a couple of files on 'update'.

                  # The user could change the file without changing its size

        Siddharth Agarwal
    
pack_dirstate: only invalidate mtime for files written in the last second...

              r19652
            
                  # within the same second. Invalidate the file's mtime in

        Siddharth Agarwal
    
dirstate: move pure python dirstate packing to pure/parsers.py

              r18567
            
                  # dirstate, forcing future 'status' calls to compare the

        Siddharth Agarwal
    
pack_dirstate: only invalidate mtime for files written in the last second...

              r19652
            
                  # contents of the file if the size is the same. This prevents

                  # mistakenly treating such files as clean.

        Siddharth Agarwal
    
parsers: inline fields of dirstate values in C version...

              r21809
            
                  e = dirstatetuple(e[0], e[1], e[2], -1)

        Siddharth Agarwal
    
dirstate: move pure python dirstate packing to pure/parsers.py

              r18567
            
                  dmap[f] = e

              if f in copymap:

                  f = "%s\0%s" % (f, copymap[f])

              e = _pack(">cllll", e[0], e[1], e[2], e[3], len(f))

              write(e)

              write(f)

          return cs.getvalue()

	Site-wide shortcuts
/	Use quick search box
g h	Goto home page
g g	Goto my private gists page
g G	Goto my public gists page
g 0-9	Goto bookmarked items from 0-9
n r	New repository page
n g	New gist page

	Repositories
g s	Goto summary page
g c	Goto changelog page
g f	Goto files page
g F	Goto files page with file search activated
g p	Goto pull requests page
g o	Goto repository settings
g O	Goto repository access permissions settings
t s	Toggle sidebar on some pages

Martin Geisler pure Python implementation of parsers.c	r7700	# parsers.py - Python implementation of parsers.c
		#
		# Copyright 2009 Matt Mackall <mpm@selenic.com> and others
		#
Martin Geisler updated license to be explicit about GPL version 2	r8225	# This software may be used and distributed according to the terms of the
Matt Mackall Update license to GPLv2+	r10263	# GNU General Public License version 2 or any later version.
Martin Geisler pure Python implementation of parsers.c	r7700
Gregory Szorc parsers: use absolute_import	r27339	from __future__ import absolute_import

		import struct
		import zlib

Yuya Nishihara parsers: switch to policy importer...	r32372	from ..node import nullid
		from .. import pycompat
Gregory Szorc util: prefer "bytesio" to "stringio"...	r36976	stringio = pycompat.bytesio
Martin Geisler pure Python implementation of parsers.c	r7700
Pulkit Goyal parsers: alias long to int on Python 3	r31220
Martin Geisler pure Python implementation of parsers.c	r7700	_pack = struct.pack
		_unpack = struct.unpack
		_compress = zlib.compress
		_decompress = zlib.decompress

Siddharth Agarwal parsers: inline fields of dirstate values in C version...	r21809	# Some code below makes tuples directly because it's more convenient. However,
		# code outside this module should always use dirstatetuple.
		def dirstatetuple(*x):
		# x is a tuple
		return x

Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133	indexformatng = ">Qiiiiii20s12x"
		indexfirst = struct.calcsize('Q')
		sizeint = struct.calcsize('i')
		indexsize = struct.calcsize(indexformatng)

		def gettype(q):
		return int(q & 0xFFFF)
Matt Mackall pure/parsers: fix circular imports, import mercurial modules properly	r7945
Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133	def offset_type(offset, type):
Martin von Zweigbergk pure: use int instead of long...	r31529	return int(int(offset) << 16 \| type)
Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133
		class BaseIndexObject(object):
		def __len__(self):
		return self._lgt + len(self._extra) + 1

		def insert(self, i, tup):
		assert i == -1
		self._extra.append(tup)
Matt Mackall pure/parsers: fix circular imports, import mercurial modules properly	r7945
Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133	def _fix_index(self, i):
		if not isinstance(i, int):
		raise TypeError("expecting int indexes")
		if i < 0:
		i = len(self) + i
		if i < 0 or i >= len(self):
		raise IndexError
		return i
Matt Mackall pure/parsers: fix circular imports, import mercurial modules properly	r7945
Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133	def __getitem__(self, i):
		i = self._fix_index(i)
		if i == len(self) - 1:
		return (0, 0, 0, -1, -1, -1, -1, nullid)
		if i >= self._lgt:
		return self._extra[i - self._lgt]
		index = self._calculate_index(i)
		r = struct.unpack(indexformatng, self._data[index:index + indexsize])
		if i == 0:
		e = list(r)
		type = gettype(e[0])
		e[0] = offset_type(0, type)
		return tuple(e)
		return r

		class IndexObject(BaseIndexObject):
		def __init__(self, data):
		assert len(data) % indexsize == 0
		self._data = data
		self._lgt = len(data) // indexsize
		self._extra = []

		def _calculate_index(self, i):
		return i * indexsize
Matt Mackall revlog: remove lazy index	r13253
Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133	def __delitem__(self, i):
Alex Gaynor style: always use `x is not None` instead of `not x is None`...	r34332	if not isinstance(i, slice) or not i.stop == -1 or i.step is not None:
Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133	raise ValueError("deleting slices only supports a:-1 with step 1")
		i = self._fix_index(i.start)
		if i < self._lgt:
		self._data = self._data[:i * indexsize]
		self._lgt = i
		self._extra = []
		else:
		self._extra = self._extra[:i - self._lgt]

		class InlinedIndexObject(BaseIndexObject):
		def __init__(self, data, inline=0):
		self._data = data
		self._lgt = self._inline_scan(None)
		self._inline_scan(self._lgt)
		self._extra = []
Martin Geisler pure Python implementation of parsers.c	r7700
Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133	def _inline_scan(self, lgt):
		off = 0
		if lgt is not None:
		self._offsets = [0] * lgt
		count = 0
		while off <= len(self._data) - indexsize:
		s, = struct.unpack('>i',
		self._data[off + indexfirst:off + sizeint + indexfirst])
		if lgt is not None:
		self._offsets[count] = off
		count += 1
		off += indexsize + s
		if off != len(self._data):
		raise ValueError("corrupted data")
		return count
Augie Fackler pure parsers: properly detect corrupt index files...	r14421
Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133	def __delitem__(self, i):
Alex Gaynor style: always use `x is not None` instead of `not x is None`...	r34332	if not isinstance(i, slice) or not i.stop == -1 or i.step is not None:
Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133	raise ValueError("deleting slices only supports a:-1 with step 1")
		i = self._fix_index(i.start)
		if i < self._lgt:
		self._offsets = self._offsets[:i]
		self._lgt = i
		self._extra = []
		else:
		self._extra = self._extra[:i - self._lgt]
Martin Geisler pure Python implementation of parsers.c	r7700
Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133	def _calculate_index(self, i):
		return self._offsets[i]
Martin Geisler pure Python implementation of parsers.c	r7700
Maciej Fijalkowski pure: write a really lazy version of pure indexObject...	r29133	def parse_index2(data, inline):
		if not inline:
		return IndexObject(data), None
		return InlinedIndexObject(data, inline), (0, data)
Martin Geisler pure Python implementation of parsers.c	r7700
		def parse_dirstate(dmap, copymap, st):
		parents = [st[:20], st[20: 40]]
Mads Kiilerich fix wording and not-completely-trivial spelling errors and bad docstrings	r17425	# dereference fields so they will be local in loop
Matt Mackall pure/parsers: fix circular imports, import mercurial modules properly	r7945	format = ">cllll"
		e_size = struct.calcsize(format)
Martin Geisler pure Python implementation of parsers.c	r7700	pos1 = 40
		l = len(st)

		# the inner loop
		while pos1 < l:
		pos2 = pos1 + e_size
		e = _unpack(">cllll", st[pos1:pos2]) # a literal here is faster
		pos1 = pos2 + e[4]
		f = st[pos2:pos1]
		if '\0' in f:
		f, c = f.split('\0')
		copymap[f] = c
		dmap[f] = e[:4]
		return parents
Siddharth Agarwal dirstate: move pure python dirstate packing to pure/parsers.py	r18567
		def pack_dirstate(dmap, copymap, pl, now):
		now = int(now)
timeless pycompat: switch to util.stringio for py3 compat	r28861	cs = stringio()
Siddharth Agarwal dirstate: move pure python dirstate packing to pure/parsers.py	r18567	write = cs.write
		write("".join(pl))
		for f, e in dmap.iteritems():
		if e[0] == 'n' and e[3] == now:
		# The file was last modified "simultaneously" with the current
		# write to dirstate (i.e. within the same second for file-
		# systems with a granularity of 1 sec). This commonly happens
		# for at least a couple of files on 'update'.
		# The user could change the file without changing its size
Siddharth Agarwal pack_dirstate: only invalidate mtime for files written in the last second...	r19652	# within the same second. Invalidate the file's mtime in
Siddharth Agarwal dirstate: move pure python dirstate packing to pure/parsers.py	r18567	# dirstate, forcing future 'status' calls to compare the
Siddharth Agarwal pack_dirstate: only invalidate mtime for files written in the last second...	r19652	# contents of the file if the size is the same. This prevents
		# mistakenly treating such files as clean.
Siddharth Agarwal parsers: inline fields of dirstate values in C version...	r21809	e = dirstatetuple(e[0], e[1], e[2], -1)
Siddharth Agarwal dirstate: move pure python dirstate packing to pure/parsers.py	r18567	dmap[f] = e

		if f in copymap:
		f = "%s\0%s" % (f, copymap[f])
		e = _pack(">cllll", e[0], e[1], e[2], e[3], len(f))
		write(e)
		write(f)
		return cs.getvalue()