upstream/mercurial-mirror Files · mercurial/logexchange.py

xdiff: add a preprocessing step that trims files...

xdiff: add a preprocessing step that trims files xdiff has a `xdl_trim_ends` step that removes common lines, unmatchable lines. That is in theory good, but happens too late - after splitting, hashing, and adjusting the hash values so they are unique. Those splitting, hashing and adjusting hash values steps could have noticeable overhead. Diffing two large files with minor (one-line-ish) changes are not uncommon. In that case, the raw performance of those preparation steps seriously matter. Even allocating an O(N) array and storing line offsets to it is expensive. Therefore my previous attempts [1] [2] cannot be good enough since they do not remove the O(N) array assignment. This patch adds a preprocessing step - `xdl_trim_files` that runs before other preprocessing steps. It counts common prefix and suffix and lines in them (needed for displaying line number), without doing anything else. Testing with a crafted large (169MB) file, with minor change: ``` open('a','w').write(''.join('%s\n' % (i % 100000) for i in xrange(30000000) if i != 6000000)) open('b','w').write(''.join('%s\n' % (i % 100000) for i in xrange(30000000) if i != 6003000)) ``` Running xdiff by a simple binary [3], this patch improves the xdiff perf by more than 10x for the above case: ``` # xdiff before this patch 2.41s user 1.13s system 98% cpu 3.592 total # xdiff after this patch 0.14s user 0.16s system 98% cpu 0.309 total # gnu diffutils 0.12s user 0.15s system 98% cpu 0.272 total # (best of 20 runs) ``` It's still slightly slower than GNU diffutils. But it's pretty close now. Testing with real repo data: For the whole repo, this patch makes xdiff 25% faster: ``` # hg perfbdiff --count 100 --alldata -c --blocks [--xdiff] # xdiff, after ! wall 0.058861 comb 0.050000 user 0.050000 sys 0.000000 (best of 100) # xdiff, before ! wall 0.077816 comb 0.080000 user 0.080000 sys 0.000000 (best of 91) # bdiff ! wall 0.117473 comb 0.120000 user 0.120000 sys 0.000000 (best of 67) ``` For files that are long (ex. commands.py), the speedup is more than 3x, very significant: ``` # hg perfbdiff --count 3000 --blocks commands.py.i 1 [--xdiff] # xdiff, after ! wall 0.690583 comb 0.690000 user 0.690000 sys 0.000000 (best of 12) # xdiff, before ! wall 2.240361 comb 2.210000 user 2.210000 sys 0.000000 (best of 4) # bdiff ! wall 2.469852 comb 2.440000 user 2.440000 sys 0.000000 (best of 4) ``` [1]: https://phab.mercurial-scm.org/D2631 [2]: https://phab.mercurial-scm.org/D2634 [3]: ``` // Code to run xdiff from command line. No proper error handling. #include <stdlib.h> #include <unistd.h> #include <sys/types.h> #include <sys/stat.h> #include <fcntl.h> #include "mercurial/thirdparty/xdiff/xdiff.h" #define ensure(x) if (!(x)) exit(255); mmfile_t readfile(const char *path) { struct stat st; int fd = open(path, O_RDONLY); fstat(fd, &st); mmfile_t file = { malloc(st.st_size), st.st_size }; ensure(read(fd, file.ptr, st.st_size) == st.st_size); close(fd); return file; } int main(int argc, char const *argv[]) { mmfile_t a = readfile(argv[1]), b = readfile(argv[2]); xpparam_t xpp = {0}; xdemitconf_t xecfg = {0}; xdemitcb_t ecb = {0}; xdl_diff(&a, &b, &xpp, &xecfg, &ecb); return 0; } ``` Differential Revision: https://phab.mercurial-scm.org/D2686

Pulkit Goyal - - Load All Authors

File last commit:

r36076:62a428bf default


                r36838:f33a87cf

default

Download file

             logexchange.py
        
                    143 lines
            
             | 4.4 KiB
            
                | text/x-python
            
             |
                PythonLexer
            
             / mercurial / logexchange.py
          
                    History
                
                 |
                  Annotation
                 | Raw
                 |Copy content
                 |Copy permalink

      # logexchange.py

      #

      # Copyright 2017 Augie Fackler <raf@durin42.com>

      # Copyright 2017 Sean Farley <sean@farley.io>

      #

      # This software may be used and distributed according to the terms of the

      # GNU General Public License version 2 or any later version.

      from __future__ import absolute_import

      from .node import hex

      from . import (

          util,

          vfs as vfsmod,

      )

      # directory name in .hg/ in which remotenames files will be present

      remotenamedir = 'logexchange'

      def readremotenamefile(repo, filename):

          """

          reads a file from .hg/logexchange/ directory and yields it's content

          filename: the file to be read

          yield a tuple (node, remotepath, name)

          """

          vfs = vfsmod.vfs(repo.vfs.join(remotenamedir))

          if not vfs.exists(filename):

              return

          f = vfs(filename)

          lineno = 0

          for line in f:

              line = line.strip()

              if not line:

                  continue

              # contains the version number

              if lineno == 0:

                  lineno += 1

              try:

                  node, remote, rname = line.split('\0')

                  yield node, remote, rname

              except ValueError:

                  pass

          f.close()

      def readremotenames(repo):

          """

          read the details about the remotenames stored in .hg/logexchange/ and

          yields a tuple (node, remotepath, name). It does not yields information

          about whether an entry yielded is branch or bookmark. To get that

          information, call the respective functions.

          """

          for bmentry in readremotenamefile(repo, 'bookmarks'):

              yield bmentry

          for branchentry in readremotenamefile(repo, 'branches'):

              yield branchentry

      def writeremotenamefile(repo, remotepath, names, nametype):

          vfs = vfsmod.vfs(repo.vfs.join(remotenamedir))

          f = vfs(nametype, 'w', atomictemp=True)

          # write the storage version info on top of file

          # version '0' represents the very initial version of the storage format

          f.write('0\n\n')

          olddata = set(readremotenamefile(repo, nametype))

          # re-save the data from a different remote than this one.

          for node, oldpath, rname in sorted(olddata):

              if oldpath != remotepath:

                  f.write('%s\0%s\0%s\n' % (node, oldpath, rname))

          for name, node in sorted(names.iteritems()):

              if nametype == "branches":

                  for n in node:

                      f.write('%s\0%s\0%s\n' % (n, remotepath, name))

              elif nametype == "bookmarks":

                  if node:

                      f.write('%s\0%s\0%s\n' % (node, remotepath, name))

          f.close()

      def saveremotenames(repo, remotepath, branches=None, bookmarks=None):

          """

          save remotenames i.e. remotebookmarks and remotebranches in their

          respective files under ".hg/logexchange/" directory.

          """

          wlock = repo.wlock()

          try:

              if bookmarks:

                  writeremotenamefile(repo, remotepath, bookmarks, 'bookmarks')

              if branches:

                  writeremotenamefile(repo, remotepath, branches, 'branches')

          finally:

              wlock.release()

      def activepath(repo, remote):

          """returns remote path"""

          local = None

          # is the remote a local peer

          local = remote.local()

          # determine the remote path from the repo, if possible; else just

          # use the string given to us

          rpath = remote

          if local:

              rpath = remote._repo.root

          elif not isinstance(remote, str):

              rpath = remote._url

          # represent the remotepath with user defined path name if exists

          for path, url in repo.ui.configitems('paths'):

              # remove auth info from user defined url

              url = util.removeauth(url)

              if url == rpath:

                  rpath = path

                  break

          return rpath

      def pullremotenames(localrepo, remoterepo):

          """

          pulls bookmarks and branches information of the remote repo during a

          pull or clone operation.

          localrepo is our local repository

          remoterepo is the peer instance

          """

          remotepath = activepath(localrepo, remoterepo)

          bookmarks = remoterepo.listkeys('bookmarks')

          # on a push, we don't want to keep obsolete heads since

          # they won't show up as heads on the next pull, so we

          # remove them here otherwise we would require the user

          # to issue a pull to refresh the storage

          bmap = {}

          repo = localrepo.unfiltered()

          for branch, nodes in remoterepo.branchmap().iteritems():

              bmap[branch] = []

              for node in nodes:

                  if node in repo and not repo[node].obsolete():

                      bmap[branch].append(hex(node))

          saveremotenames(localrepo, remotepath, bmap, bookmarks)

	Site-wide shortcuts
/	Use quick search box
g h	Goto home page
g g	Goto my private gists page
g G	Goto my public gists page
g 0-9	Goto bookmarked items from 0-9
n r	New repository page
n g	New gist page

	Repositories
g s	Goto summary page
g c	Goto changelog page
g f	Goto files page
g F	Goto files page with file search activated
g p	Goto pull requests page
g o	Goto repository settings
g O	Goto repository access permissions settings
t s	Toggle sidebar on some pages

				# logexchange.py
				#
				# Copyright 2017 Augie Fackler <raf@durin42.com>
				# Copyright 2017 Sean Farley <sean@farley.io>
				#
				# This software may be used and distributed according to the terms of the
				# GNU General Public License version 2 or any later version.

				from __future__ import absolute_import

				from .node import hex

				from . import (
				util,
				vfs as vfsmod,
				)

				# directory name in .hg/ in which remotenames files will be present
				remotenamedir = 'logexchange'

				def readremotenamefile(repo, filename):
				"""
				reads a file from .hg/logexchange/ directory and yields it's content
				filename: the file to be read
				yield a tuple (node, remotepath, name)
				"""

				vfs = vfsmod.vfs(repo.vfs.join(remotenamedir))
				if not vfs.exists(filename):
				return
				f = vfs(filename)
				lineno = 0
				for line in f:
				line = line.strip()
				if not line:
				continue
				# contains the version number
				if lineno == 0:
				lineno += 1
				try:
				node, remote, rname = line.split('\0')
				yield node, remote, rname
				except ValueError:
				pass

				f.close()

				def readremotenames(repo):
				"""
				read the details about the remotenames stored in .hg/logexchange/ and
				yields a tuple (node, remotepath, name). It does not yields information
				about whether an entry yielded is branch or bookmark. To get that
				information, call the respective functions.
				"""

				for bmentry in readremotenamefile(repo, 'bookmarks'):
				yield bmentry
				for branchentry in readremotenamefile(repo, 'branches'):
				yield branchentry

				def writeremotenamefile(repo, remotepath, names, nametype):
				vfs = vfsmod.vfs(repo.vfs.join(remotenamedir))
				f = vfs(nametype, 'w', atomictemp=True)
				# write the storage version info on top of file
				# version '0' represents the very initial version of the storage format
				f.write('0\n\n')

				olddata = set(readremotenamefile(repo, nametype))
				# re-save the data from a different remote than this one.
				for node, oldpath, rname in sorted(olddata):
				if oldpath != remotepath:
				f.write('%s\0%s\0%s\n' % (node, oldpath, rname))

				for name, node in sorted(names.iteritems()):
				if nametype == "branches":
				for n in node:
				f.write('%s\0%s\0%s\n' % (n, remotepath, name))
				elif nametype == "bookmarks":
				if node:
				f.write('%s\0%s\0%s\n' % (node, remotepath, name))

				f.close()

				def saveremotenames(repo, remotepath, branches=None, bookmarks=None):
				"""
				save remotenames i.e. remotebookmarks and remotebranches in their
				respective files under ".hg/logexchange/" directory.
				"""
				wlock = repo.wlock()
				try:
				if bookmarks:
				writeremotenamefile(repo, remotepath, bookmarks, 'bookmarks')
				if branches:
				writeremotenamefile(repo, remotepath, branches, 'branches')
				finally:
				wlock.release()

				def activepath(repo, remote):
				"""returns remote path"""
				local = None
				# is the remote a local peer
				local = remote.local()

				# determine the remote path from the repo, if possible; else just
				# use the string given to us
				rpath = remote
				if local:
				rpath = remote._repo.root
				elif not isinstance(remote, str):
				rpath = remote._url

				# represent the remotepath with user defined path name if exists
				for path, url in repo.ui.configitems('paths'):
				# remove auth info from user defined url
				url = util.removeauth(url)
				if url == rpath:
				rpath = path
				break

				return rpath

				def pullremotenames(localrepo, remoterepo):
				"""
				pulls bookmarks and branches information of the remote repo during a
				pull or clone operation.
				localrepo is our local repository
				remoterepo is the peer instance
				"""
				remotepath = activepath(localrepo, remoterepo)
				bookmarks = remoterepo.listkeys('bookmarks')
				# on a push, we don't want to keep obsolete heads since
				# they won't show up as heads on the next pull, so we
				# remove them here otherwise we would require the user
				# to issue a pull to refresh the storage
				bmap = {}
				repo = localrepo.unfiltered()
				for branch, nodes in remoterepo.branchmap().iteritems():
				bmap[branch] = []
				for node in nodes:
				if node in repo and not repo[node].obsolete():
				bmap[branch].append(hex(node))

				saveremotenames(localrepo, remotepath, bmap, bookmarks)